What Is a Feature Flag? A Working Engineer's Guide

Summary

A feature flag is a conditional in your code that turns behavior on or off at runtime, without a redeploy. It separates deploying code from releasing it, so you can merge unfinished work, roll out gradually, and kill a feature fast. Flags come in four types: release, experiment, ops, and permission. Each flag is also a fork in your code, so every temporary one needs an owner, an expiry date, and a removal ticket from day one.

Engineer desk with a row of physical toggle switches beside a laptop

A feature flag is a conditional in your code that decides at runtime whether a piece of behavior is on or off, without a new deploy. That is the whole idea. What is a feature flag in practice? An if statement whose answer comes from config, a database row, or a flag service instead of being hardcoded.

The concept is simple. Living with 200 of them in a repo is not, and that second part is what this guide is really about.

What does a feature flag look like in code?

Here is the smallest useful version:

if (flags.isEnabled("new-checkout", { userId })) {
  return renderNewCheckout();
}
return renderOldCheckout();

Both code paths ship in the same build. The flag value lives somewhere you can change without redeploying: an environment variable, a JSON file, a table, or a hosted service. Change the value, and behavior changes on the next evaluation.

That separation is the point. Deployment is putting code on servers. Release is letting users see it. Flags split those two events, so merging to main no longer means "everyone gets this now."

Hand flipping a single toggle switch with a green indicator light

Why do teams use feature flags at all?

Three reasons come up again and again in real teams of 5 to 50 engineers.

None of these require a vendor. A flag can be a boolean in a config table. Skip the platform until you need percentage rollouts, audit trails, or non-engineers toggling things.

What are the four types of feature flags?

Pete Hodgson's widely cited feature toggles article on Martin Fowler's site sorts flags by how long they live and how often the decision changes. The categories are still the clearest mental model.

The lifespan column matters most. A release flag that outlives its release is a bug you have not found yet. A permission flag that gets deleted in a cleanup sprint is an outage you have not had yet.

Name and label the type when you create the flag. Six months later, nobody remembers which category "new-nav-v2" was.

Team planning grid of sticky notes grouped by category

How do you roll out a flag without hurting users?

The boring sequence works best.

  1. Ship the code with the flag off. Verify nothing changed.

  2. Enable for your own team in production. Use it for a day.

  3. Enable for 1% to 5% of users, sticky by user ID so nobody flips between variants mid-session.

  4. Watch error rate, latency, and one business metric. Decide the threshold before you start, not while staring at a dashboard.

  5. Ramp to 100%, wait a defined period, then remove the flag.

Step five is the one teams skip. We will get to why that costs more than you think.

A practical note on evaluation: make flag reads cheap and failure-safe. If the flag service is down, your code needs a default. Pick the default per flag, on purpose. A kill switch should fail to "safe," while a new feature should fail to "off."

Why do feature flags become technical debt?

Every flag is a fork in your code. Two flags make four possible paths, ten flags make 1,024, and you almost certainly tested a handful of them. GrowthBook's engineering guide to flag debt cites research showing that about 75% of toggle components were still in codebases up to 49 weeks after being introduced, even though most developers said they planned to remove them.

Hodgson puts it well in the same Fowler article: savvy teams treat toggles as inventory with a carrying cost, and work to keep that inventory low.

The classic horror story is Knight Capital in 2012. A retired feature's flag was reused for new behavior while old code still sat on one server. That mismatch contributed to roughly $460 million in losses in under an hour. Your stale flag probably will not do that. It will do something quieter: a refactor that breaks a branch nobody knew existed, or a new hire spending an afternoon figuring out which of two checkout paths is real.

Dusty shelf of forgotten boxes and old switches

How do you find every flag in an existing codebase?

This is the question the vendor docs skip, and it is where most teams get stuck. The flag dashboard tells you what is configured. It does not tell you where each flag is read in code, or whether the code path behind an "off" flag is still reachable.

Start with the cheap approach:

# every literal flag key read via your wrapper
rg -n 'isEnabled\("' src/ | sort

# flags defined but never referenced
comm -23 <(jq -r 'keys[]' flags.json | sort) \
         <(rg -o 'isEnabled\("([^"]+)"' -r '$1' src/ | sort -u)

This breaks down fast. Flag keys built from string concatenation do not show up in grep. Flags passed through helper functions hide their call sites. Monorepos and multi-repo setups multiply the problem, because the same key may be read by three services.

That is where reading the code with tooling helps. Code search tools that understand your repo can answer "where is new-checkout evaluated, and what depends on the result?" in one query instead of an afternoon of grep. Cursor and GitHub Copilot handle this reasonably inside one repo. Across several repos, you need an index that spans all of them, which is the case we built codebasechat for.

What does a sane flag cleanup process look like?

Treat removal as part of the work, not a chore for later.

Static analysis helps here as well. A tool like CodeScene can show which files carry the most tangled conditional logic, which is usually where old flags cluster. SonarQube flags unreachable and dead code after you remove a check.

How do you test code that sits behind a flag?

Testing is the part nobody budgets for. With two paths per flag, your test suite needs to cover both, at least for the flags that guard risky behavior.

Keep it practical. Test the on and off states of every release flag in unit tests, by injecting the flag value rather than reading a live service. Run one end-to-end suite against the default production configuration, because that is what users get today. Then run a second pass with the flag on for the feature you are about to release.

Do not try to test every combination. With ten flags you cannot. Instead, keep flags independent: a flag that changes behavior only when another flag is also on is a design smell, and it is the first thing to untangle.

What do new engineers get wrong about flags?

Juniors tend to make the same three mistakes, and each is cheap to prevent in review.

The first two weeks of a new hire are exactly when they run into old flags with no owner. A short flag registry, with a type, an owner, and a removal date for each entry, saves several hours of asking around.

When should you skip feature flags?

Flags are not free, so skip them when:

And use them without hesitation when a change is risky, user-facing, and hard to reverse by redeploying. Payment flows, auth changes, and anything with a data migration behind it all qualify.

What we would actually do on a team of ten

Start with a boolean in config and one wrapper function, so every flag read goes through one place. That single choke point makes the flag list greppable, auditable, and easy to migrate to a hosted service later.

Label each flag by type, give it an owner, and file the removal ticket on day one. Review the list monthly for ten minutes. Delete the flags that have been at 100% for two weeks.

A feature flag is a loan. Take it when it saves you a risky release, and pay it back before the interest shows up in your next refactor.

Frequently asked questions

What is a feature flag in simple terms?
A feature flag is an if statement whose answer comes from config or a flag service instead of being hardcoded. It lets you turn a feature on or off for some or all users without deploying new code.
What is the difference between a feature flag and a feature toggle?
Nothing meaningful. Feature flag, feature toggle, and feature switch describe the same technique. Different teams and vendors simply prefer different words.
What are the main types of feature flags?
Release flags hide unfinished work, experiment flags power A/B tests, ops flags act as kill switches or load controls, and permission flags gate features by user group. They differ mostly in how long they should live.
Do you need a paid tool to use feature flags?
No. A boolean in a config table plus one wrapper function is enough to start. Move to a hosted service when you need percentage rollouts, audit logs, or non-engineers toggling flags.
How long should a feature flag live?
Release and experiment flags should be removed within weeks of reaching 100% or a decision. Around 90 days without a change is a common review trigger. Permission flags and true kill switches can live indefinitely.
How do you find stale feature flags in a codebase?
Compare the flags configured in your dashboard against the keys actually read in code, then check which have served 100% of traffic for two weeks or more. Grep works for literal keys, but dynamic keys and multi-repo setups need code search that understands the whole repo.
Are feature flags bad for code quality?
Only when they are never removed. Each flag adds a code path and doubles the states you should test in principle, so unmanaged flags become technical debt. Owners, expiry dates, and removal tickets keep them under control.