App Store A/B Testing in the Wild: What 53,000 Product Page Experiments Reveal

By AppSigma Team · Published July 22, 2026

TL;DR: App Store A/B testing is no longer a power-user tactic — it's the default behavior of every app with something to lose. Over the last 11 months our crawlers detected 53,354 live Product Page Optimization experiments across 23,893 iOS apps. Screenshots are what everyone tests (99% of experiments), the median test runs 11 days, and the apps at the top of the store test the hardest: 62% of today's US top-grossing top 100 have run at least one experiment. Roughly 5,000 new tests go live every month — and your competitors' tests are visible, if you know how to look.

The numbers below come from our own detection pipeline: when Apple serves a product-page variant to one of our crawl sessions, we record the experiment, its variants, and its creative assets — the same infrastructure behind our App Store API and MCP server. Observation window: August 23, 2025 – July 22, 2026.

How we see "invisible" experiments

Apple's Product Page Optimization (PPO) lets developers run up to three treatment variants against their default product page. Apple never publishes who is testing what — results live privately in App Store Connect. But the tests themselves run in public: a share of real store traffic is served the variant pages. When our sessions land in a treatment bucket, the experiment reveals itself — a different screenshot set, a different icon, different promotional text, served under the same app listing.

Each detected experiment is confirmed across multiple observations, with its variant creatives hashed and archived. One observation is a curiosity; 53,354 experiments in 11 months is a census.

What the App Store actually tests

What varies Share of experiments
Screenshots 99% (52,871 of 53,354)
Promotional text 28% (14,846)
App icon 7% (3,856)

Screenshots are effectively the unit of App Store conversion testing. Icon tests — the thing most "ASO checklist" articles obsess over — barely register at 7%, and they usually ride along inside a broader screenshot test rather than running alone.

Variant counts split cleanly: 60% of experiments run a single treatment against control, 17% run three variants, and 22% run the full four (control plus the maximum three treatments). Running four variants costs traffic — each arm gets a quarter of impressions — so the four-variant crowd skews toward high-traffic apps that can afford the split.

How long does a test run?

The median experiment lasts 11 days from first to last observation. A quarter are done within 5 days, a quarter run past 29 days, and one in ten stretches past 67 — those long tails are mostly tests someone forgot to stop, still serving variants months later.

That 11-day median is worth comparing to Apple's own ceiling of 90 days: in practice, apps that test seriously call their experiments in under two weeks and move to the next one.

And they do move to the next one. Among apps with five or more experiments, the median gap between one test ending and the next beginning is under 8 days, and 11% of follow-up tests start within 48 hours. For the serious testers, PPO isn't an event — it's a pipeline.

Who tests: the top of the store

We checked today's US top-100 charts against the experiment log:

Chart Apps that ran ≥1 test in 11 months
Top grossing top 100 62%
Free top 100 48%
Paid top 100 6%

The pattern is blunt: the more money a store page makes, the more likely it is being tested right now. Paid apps — a shrinking, conversion-insensitive corner of the store — barely test at all. And at any given moment, about 2,900 experiments are live on the store.

Most apps that test at all test once: of the 23,893 apps in our log, 14,440 ran exactly one experiment. But 2,495 apps ran five or more, and 641 ran ten or more — a dedicated testing class that treats the product page as a permanent experiment surface.

The heaviest testers on the App Store

The top of the leaderboard is not who you'd expect. Not the casino giants, not the big social apps — utility and AI apps:

App Experiments in 11 months
Greatnotes5: AI Meeting Notes 49
AI Home Design: Room Planner 41
Remove Objects・AI Photo Eraser 39
Photo Collage Maker・Pic Layout 35
Blur Photo - Background & Face 32
PDF Converter – Photos to PDF 32

Forty-nine experiments in eleven months is a new test roughly every 8 days, continuously, for a year. These are subscription utilities living and dying on install-to-trial conversion, where a percentage point of screenshot conversion compounds across every keyword they rank for. (Casino apps are close behind — the heaviest slots publishers ran 15–20 tests each.)

We've also seen what disciplined testing looks like when it's part of a bigger play: in our Keiki case study, a kids-learning app ran nine screenshot experiments, settled its winning creative, and then shipped a hidden keyword-field expansion into a page already tuned to convert — and held the rankings it gained.

Why this matters for your ASO

  1. Testing is table stakes at the top. If you're competing against top-grossing apps, assume their page has been through a dozen experiments yours hasn't.
  2. Test screenshots first. The market has voted: 99% of experiments touch screenshots. Your icon is probably fine.
  3. Two weeks per test is the working rhythm. The median real-world test is 11 days. Budget traffic accordingly, and don't let a test idle for months.
  4. Conversion work protects ranking work. Rankings you can't convert are rankings you'll lose — keyword positions are re-scored continuously, and post-tap conversion feeds the score.
  5. Your competitors' tests are watchable. A test on a competitor's page tells you exactly which message they're trying to win with — before their new creative even becomes the default.

If you want to know when a competitor starts testing — or what the rest of your category is converging on — AppSigma tracks product pages, rankings and creatives across the store daily.

Method notes: an "experiment" is a distinct PPO test detected on a distinct app and locale; 90% ran on the US storefront. Durations are first-to-last observation, so tests shorter than our crawl interval can be missed and our counts undercount reality. We observe which variants exist and when they run — not their conversion results, which stay private to the developer.

Frequently asked questions

What is App Store Product Page Optimization?

Product Page Optimization (PPO) is Apple's built-in A/B testing tool in App Store Connect. It lets developers test up to three treatment versions of their product page — different screenshots, icon, or promotional text — against the default page on a share of real App Store traffic.

How long should an App Store A/B test run?

Apple allows up to 90 days, but real-world behavior is much faster: across 53,354 detected experiments, the median test ran about 11 days, and serious testers started their next experiment within 8 days of ending the last one.

What do apps test most on their product pages?

Screenshots, overwhelmingly — 99% of the experiments we detected varied screenshots, 28% varied promotional text, and only 7% touched the app icon. Screenshot creative is where App Store conversion testing actually happens.

Can you see competitors' App Store A/B tests?

Not through App Store Connect — results are private. But the tests themselves run on public store traffic, so variant pages can be detected by monitoring product pages at scale. AppSigma detects live experiments, their variant creatives, and when they start and end.