Ship behind a flag
decide on real numbers, not on who argued longer
Shipping a change to everyone at once and hoping it helps is still how most teams ship. We build feature flag infrastructure so a change goes to a small slice of users first, gets measured against a metric that matters, and rolls back in seconds if it does not help.
What it is
A feature flag wraps a piece of code in a condition: show the new checkout flow to this percentage of users, or to users with this attribute, and the old flow to everyone else, controlled by a switch that does not require a deploy to change. Experiment infrastructure builds on top of that, assigning users consistently to a test group, measuring a defined metric for each group, and giving a statistically grounded answer to whether the change actually helped.
When you need it (and when you do not)
You need this once a change is risky enough that rolling it out to everyone at once is a real gamble, a new pricing page, a redesigned onboarding flow, a changed algorithm, where being wrong costs real revenue or real users. It is also the right build once “ship it and see” has produced a few disagreements about whether a past change actually helped, which usually means nobody measured it properly at the time.
You do not need this for low-risk changes, a copy tweak, a color change, a bug fix, where the cost of being wrong is trivial and a flag adds process without adding value. The tell that you need it is a change where someone in the room says “I think this will help” and someone else disagrees, and neither has data to settle it.
How we build it
For most clients already on PostHog for product analytics, we use its built-in feature flags, which keeps assignment, tracking and results in one tool rather than stitching two systems together. Rollout rules support both simple percentage-based assignment and attribute-based targeting, a specific plan tier or region, for cases where a blanket percentage is not the right test. Every flag has an instant kill switch, flipping it off takes effect immediately with no deploy, which is the entire point: a bad change should be reversible in seconds, not after an incident review and a hotfix. For experiments specifically, we help define the success metric and minimum sample size before launch, not after results come in, because a metric chosen after the fact tends to be the one that makes the result look good rather than the one that actually answers the question.
What to watch
The real risk with feature flags is not technical, it is organizational drift: flags that were meant to be temporary for a rollout stay in the codebase for years, and eventually nobody remembers which combination of flags is actually live for which users, turning the code itself into a maze of conditionals. We recommend a flag cleanup review on a schedule, and build flags to be removed once a rollout completes, not left in indefinitely by default. On the experimentation side, calling a result significant before the sample size is reached is the most common statistical mistake teams make on their own; we build a minimum sample size check into the reporting so a test does not get called early just because the numbers looked good on day two.
Sample ratio mismatch is a subtler failure worth watching for: if your assignment logic is even slightly biased, say by excluding users on an older app version from one group but not the other, the groups are no longer comparable and the result is not trustworthy no matter how significant it looks. We check for this before reading any result as final, since it is a common and easy-to-miss way an otherwise correct-looking experiment produces a wrong answer that nobody catches until a decision has already been made on it.
Price and timeline
| Scope | Price | Timeline |
|---|---|---|
| Flags, rollout, kill switch | from $1,200 | 1 to 2 weeks |
| Full experimentation with significance testing | from $2,800 | 2 to 3 weeks |
Related
Built as part of custom development and analytics. Builds directly on product analytics setup for measuring results. See the funnel work behind the trading app funnel audit and the engagement mechanics in an AI coach and gamification fitness app. Tell us what change you want to test before committing to it: get in touch.
FAQ
How much does feature flag infrastructure cost?
A core setup with percentage rollout and a kill switch starts at $1,200. A fuller experimentation platform with statistical significance testing and analytics integration runs $2,500 to $4,500.
How long does it take?
1 to 3 weeks depending on whether flags need to work across both frontend and backend, and how many existing features need to be wrapped in a flag versus just new ones going forward.
Do we need PostHog or a dedicated tool?
PostHog's built-in feature flags cover most needs well and pair naturally with its analytics. A dedicated tool like LaunchDarkly makes sense at a scale, or compliance requirement, that most of our clients have not reached yet, we say so rather than oversell it.
What is the actual benefit over just deploying changes?
A bad deploy without a flag means a rollback, a new deploy, and downtime for everyone in between. A bad change behind a flag means flipping a switch, no deploy, no downtime, while only the test group was ever affected.
Who decides what counts as a successful experiment?
You do, before the experiment starts. We help define a clear success metric and a minimum sample size upfront specifically so the result is not re-interpreted after the fact to match whatever answer someone wanted.