What Is a Feature Flag? A Working Engineer's Guide
Summary
A feature flag is a conditional in your code that turns behavior on or off at runtime, without a redeploy. It separates deploying code from releasing it, so you can merge unfinished work, roll out gradually, and kill a feature fast. Flags come in four types: release, experiment, ops, and permission. Each flag is also a fork in your code, so every temporary one needs an owner, an expiry date, and a removal ticket from day one.
A feature flag is a conditional in your code that decides at runtime whether a piece of behavior is on or off, without a new deploy. That is the whole idea. What is a feature flag in practice? An if statement whose answer comes from config, a database row, or a flag service instead of being hardcoded.
The concept is simple. Living with 200 of them in a repo is not, and that second part is what this guide is really about.
What does a feature flag look like in code?
Here is the smallest useful version:
if (flags.isEnabled("new-checkout", { userId })) {
return renderNewCheckout();
}
return renderOldCheckout();Both code paths ship in the same build. The flag value lives somewhere you can change without redeploying: an environment variable, a JSON file, a table, or a hosted service. Change the value, and behavior changes on the next evaluation.
That separation is the point. Deployment is putting code on servers. Release is letting users see it. Flags split those two events, so merging to main no longer means "everyone gets this now."

Why do teams use feature flags at all?
Three reasons come up again and again in real teams of 5 to 50 engineers.
Merge unfinished work safely. You commit the half-built feature behind a flag that is off, and your branch never lives for three weeks. This is what makes trunk-based development workable.
Roll out gradually. Turn a change on for internal users, then 5% of traffic, then everyone. If error rates climb, you flip it back in seconds.
Kill something fast. When a payment provider misbehaves at 2 a.m., a flag lets on-call degrade one feature instead of rolling back an entire release.
None of these require a vendor. A flag can be a boolean in a config table. Skip the platform until you need percentage rollouts, audit trails, or non-engineers toggling things.
What are the four types of feature flags?
Pete Hodgson's widely cited feature toggles article on Martin Fowler's site sorts flags by how long they live and how often the decision changes. The categories are still the clearest mental model.
Release: lives days to weeks, flipped by engineers. Example: hide an unfinished checkout redesign.
Experiment: lives weeks, flipped by product and data. Example: an A/B test of two pricing pages.
Ops: lives hours to forever, flipped by on-call. Example: a kill switch for a slow recommendations service.
Permission: lives months to years, flipped by product and support. Example: beta access or premium-only features.
The lifespan column matters most. A release flag that outlives its release is a bug you have not found yet. A permission flag that gets deleted in a cleanup sprint is an outage you have not had yet.
Name and label the type when you create the flag. Six months later, nobody remembers which category "new-nav-v2" was.

How do you roll out a flag without hurting users?
The boring sequence works best.
Ship the code with the flag off. Verify nothing changed.
Enable for your own team in production. Use it for a day.
Enable for 1% to 5% of users, sticky by user ID so nobody flips between variants mid-session.
Watch error rate, latency, and one business metric. Decide the threshold before you start, not while staring at a dashboard.
Ramp to 100%, wait a defined period, then remove the flag.
Step five is the one teams skip. We will get to why that costs more than you think.
A practical note on evaluation: make flag reads cheap and failure-safe. If the flag service is down, your code needs a default. Pick the default per flag, on purpose. A kill switch should fail to "safe," while a new feature should fail to "off."
Why do feature flags become technical debt?
Every flag is a fork in your code. Two flags make four possible paths, ten flags make 1,024, and you almost certainly tested a handful of them. GrowthBook's engineering guide to flag debt cites research showing that about 75% of toggle components were still in codebases up to 49 weeks after being introduced, even though most developers said they planned to remove them.
Hodgson puts it well in the same Fowler article: savvy teams treat toggles as inventory with a carrying cost, and work to keep that inventory low.
The classic horror story is Knight Capital in 2012. A retired feature's flag was reused for new behavior while old code still sat on one server. That mismatch contributed to roughly $460 million in losses in under an hour. Your stale flag probably will not do that. It will do something quieter: a refactor that breaks a branch nobody knew existed, or a new hire spending an afternoon figuring out which of two checkout paths is real.

How do you find every flag in an existing codebase?
This is the question the vendor docs skip, and it is where most teams get stuck. The flag dashboard tells you what is configured. It does not tell you where each flag is read in code, or whether the code path behind an "off" flag is still reachable.
Start with the cheap approach:
# every literal flag key read via your wrapper
rg -n 'isEnabled\("' src/ | sort
# flags defined but never referenced
comm -23 <(jq -r 'keys[]' flags.json | sort) \
<(rg -o 'isEnabled\("([^"]+)"' -r '$1' src/ | sort -u)This breaks down fast. Flag keys built from string concatenation do not show up in grep. Flags passed through helper functions hide their call sites. Monorepos and multi-repo setups multiply the problem, because the same key may be read by three services.
That is where reading the code with tooling helps. Code search tools that understand your repo can answer "where is new-checkout evaluated, and what depends on the result?" in one query instead of an afternoon of grep. Cursor and GitHub Copilot handle this reasonably inside one repo. Across several repos, you need an index that spans all of them, which is the case we built codebasechat for.
What does a sane flag cleanup process look like?
Treat removal as part of the work, not a chore for later.
Create the removal ticket with the flag. Link it in the flag description. If the ticket does not exist, the flag does not ship.
Set an owner and an expiry date on every non-permanent flag. Around 90 days without a change is a reasonable trigger for review.
Remove in two pull requests. First delete the flag check and keep the winning path. Then delete the dead branch and its tests. Small diffs are reviewable diffs.
Add a time bomb test. A test that fails when a release flag passes its expiry date turns a good intention into a red build.
Cap the total. If you have 40 active flags and the limit is 40, adding one means removing one.
Static analysis helps here as well. A tool like CodeScene can show which files carry the most tangled conditional logic, which is usually where old flags cluster. SonarQube flags unreachable and dead code after you remove a check.
How do you test code that sits behind a flag?
Testing is the part nobody budgets for. With two paths per flag, your test suite needs to cover both, at least for the flags that guard risky behavior.
Keep it practical. Test the on and off states of every release flag in unit tests, by injecting the flag value rather than reading a live service. Run one end-to-end suite against the default production configuration, because that is what users get today. Then run a second pass with the flag on for the feature you are about to release.
Do not try to test every combination. With ten flags you cannot. Instead, keep flags independent: a flag that changes behavior only when another flag is also on is a design smell, and it is the first thing to untangle.
What do new engineers get wrong about flags?
Juniors tend to make the same three mistakes, and each is cheap to prevent in review.
Nesting flags. One flag inside another creates a path that only exists when both are on. Ask for a single flag with a clear name instead.
Putting logic in the flag name. A key like "show-new-nav-to-premium-users-in-eu" encodes a targeting rule that belongs in the flag service, not in a string.
Forgetting the default. If the flag lookup fails, what happens? Make the author write the answer in the pull request description.
The first two weeks of a new hire are exactly when they run into old flags with no owner. A short flag registry, with a type, an owner, and a removal date for each entry, saves several hours of asking around.
When should you skip feature flags?
Flags are not free, so skip them when:
The change is small, reversible, and one deploy away from a rollback. A flag adds a code path for no gain.
The change touches a database schema in a way a flag cannot hide. Use expand-and-contract migrations instead.
Your team has no process for removing flags. Fix that first, or you are borrowing against future readability.
And use them without hesitation when a change is risky, user-facing, and hard to reverse by redeploying. Payment flows, auth changes, and anything with a data migration behind it all qualify.
What we would actually do on a team of ten
Start with a boolean in config and one wrapper function, so every flag read goes through one place. That single choke point makes the flag list greppable, auditable, and easy to migrate to a hosted service later.
Label each flag by type, give it an owner, and file the removal ticket on day one. Review the list monthly for ten minutes. Delete the flags that have been at 100% for two weeks.
A feature flag is a loan. Take it when it saves you a risky release, and pay it back before the interest shows up in your next refactor.