Post A got 40 likes. Post B got 52. Someone on the team declares B the winner, the team adopts B's hook for next month's content, and everyone moves on.
Nothing about that conclusion is supported by the numbers. Neither post's reach was accounted for — if A reached 500 people and B reached 900, A actually had the higher engagement rate. Neither result was checked against the kind of random variation you'd expect from posting the same content twice. And "52 beat 40" carries no information about whether that gap would hold up next week, or whether it was noise.
This is not a pedantic statistics complaint. It's the difference between a team that's actually learning what works and a team that's confidently wrong on a monthly cadence.
What makes a test valid in the first place
A real experiment needs three things a like-count comparison skips entirely: a consistent denominator, a large enough sample, and a way to tell a real difference from random noise.
If you can't answer "out of how many people did each version reach?", you don't have a test result yet — you have two anecdotes.
- A consistent denominator — compare rates, not raw counts. "40 likes" means nothing without "out of how many people saw it." Engagement rate (interactions ÷ reach) is the number that's actually comparable between two posts
- A large enough sample — a post reaching 60 people can swing from 2% to 8% engagement rate on pure chance. The smaller the reach, the less any single result tells you
- A significance check — a formal way of asking "how likely is it that this gap happened by chance alone?" before you act on it
The test in plain terms: a two-proportion z-test
Engagement rate is a proportion — interactions out of reach, the same shape as "conversions out of visitors" in any other kind of A/B test. That means the standard tool for comparing two proportions applies directly: a two-proportion z-test.
The mechanics, without the notation: it takes each variant's engagement rate and sample size, computes how far apart those two rates are relative to how much variation you'd expect from sample size alone, and converts that into a confidence level — the probability that the gap you're seeing is real rather than random.
// Conceptually — not the exact implementation, but the shape of it
function estimateConfidence(a, b) {
const rateA = a.interactions / a.reach;
const rateB = b.interactions / b.reach;
// pooled rate + standard error account for BOTH sample sizes —
// a 200-person sample and a 20,000-person sample don't get equal weight
const pooled = (a.interactions + b.interactions) / (a.reach + b.reach);
const stdError = Math.sqrt(pooled * (1 - pooled) * (1 / a.reach + 1 / b.reach));
const z = Math.abs(rateA - rateB) / stdError;
return zScoreToConfidence(z); // e.g. z ≈ 1.96 → ~95% confidence
}The two inputs that actually decide whether you have a real result: how big the gap is, and how much data backs each side of it. A 3-point gap on 10,000 people each is far more trustworthy than the same 3-point gap on 80 people each — and a raw like-count comparison has no way to tell those two situations apart.
Why a made-up confidence number is worse than none
It's tempting to skip the math and just type a number into a report: "we're about 85% confident B wins." That number feels rigorous. It isn't — and it's actively worse than admitting you don't know, because it launders a guess into something that looks like evidence.
A fabricated confidence score does three specific kinds of damage: it can't be checked by anyone else on the team, it doesn't account for sample size at all (an 85% guess "feels" the same whether it's based on 50 people or 50,000), and it trains the team to trust round, confident-sounding numbers over the harder discipline of actually running the test to completion.
A real 62% confidence result and an honest "we don't have enough data yet" are both more useful than a hand-typed 85% — because both of the first two are true.
What to actually test
The gap between a useful experiment and a wasted one is usually variable isolation — testing one thing at a time, so a winning result tells you *what* worked.
- Hook / opening line — the single highest-leverage thing to test; it decides whether someone reads the rest
- Format — question vs. statement, story vs. list, single image vs. carousel — but only one axis per test
- Call-to-action — "comment below" vs. "link in bio" vs. no explicit CTA at all
- Length — same message, materially different length, nothing else changed
Changing the hook, the image, and the CTA all in one "test" produces a result you can't act on — you'll know one version won, not why. Pick one variable, hold everything else constant, and let the sample size actually accumulate before calling it.
How Postlore approaches this
Postlore's A/B experiments run the real version of this, not the eyeball version. Publish two variants of a post, and Postlore tracks reach and engagement for both, computes a genuine two-proportion z-test confidence, and auto-concludes the moment both variants have a full 7-day analytics window — no one on the team has to remember to check back or do the math by hand.
There's no field anywhere to type in a confidence number. It's computed server-side from the same reach and engagement data the report shows, so the number on the card is always the number the math actually produced — and when there isn't yet enough data to say anything, it says so instead of guessing.
Key takeaways
- Comparing raw like counts between two posts isn't a test — compare engagement rate (interactions ÷ reach), the number that actually accounts for how many people saw each version
- Small sample sizes swing wildly on pure chance; a gap on 60 people means far less than the same gap on 6,000
- A two-proportion z-test is the standard tool for this — it weighs both the size of the gap and the sample size behind each side of it
- A hand-typed confidence percentage is worse than no number: it can't be verified, ignores sample size, and trains a team to trust false precision
- Isolate one variable per test — hook, format, CTA, or length — so a winning result tells you what to repeat, not just that something changed