All insights
Software development
•11 min read
Feature flag system design: rules you won't regret in eighteen months

Feature flag system design: rules you won't regret in eighteen months

Feature flag system design for small B2B teams: build vs buy, local evaluation, safe defaults, audit trails and expiry rules.

Author
By DevLume
Published
Published 3 October 2026

Key takeaways

  • Flags aren't one thing. Pete Hodgson's taxonomy separates release, experiment, ops and permission toggles, and each needs a different lifetime and owner (martinfowler.com, 2017).
  • Flag debt is measurable. In Chrome, 53% of release toggles were still in the code after 10 releases (Rahman et al., MSR 2016). Uber built a tool that generated cleanup diffs for 1,381 stale flags, 17% of all its flags (Ramanathan et al., ICSE-SEIP 2020).
  • Evaluate server-side flags in process, against a cached ruleset. That's how LaunchDarkly, Unleash and Flagsmith's local mode work, and it keeps the flag service out of your request path.
  • Every evaluation needs a safe default for when the flag service is unreachable. The OpenFeature spec makes this a hard rule: evaluation "must always return the default value in the event of abnormal execution" (OpenFeature).
  • The best-documented flag failure came from reusing an old flag. Knight Capital lost over $460 million in about 45 minutes (SEC, 2013).

TL;DR

A feature flag system that's still healthy eighteen months from now comes down to six decisions. First, separate flag types and give each a lifetime. Second, evaluate in process, not over the network. Third, pick a deliberate default for every flag so an outage doesn't change what users see. Fourth, log every change with who made it and why. Fifth, never reuse a flag name. Sixth, make cleanup part of shipping a feature, not a separate project. For most small and mid-sized B2B teams, buying a hosted service (or running an open-source one behind the OpenFeature SDK) beats building one. The rules matter more than the tool, though, and you'll need them either way.

Why do feature flag systems go wrong?

They fail slowly, then all at once. A flag starts as a one-line if statement that makes a risky release reversible. Nobody removes it. Six months later there are 200 of them, nobody knows which ones are live, and testing every combination of flag states stopped being possible long ago.

Pete Hodgson's long-standing guide on martinfowler.com describes the right mental model: "Savvy teams view their Feature Toggles as inventory which comes with a carrying cost, and work to keep that inventory as low as possible" (martinfowler.com, 2017). Every live flag doubles the paths through one piece of code. It also adds a config value someone can change without a deploy, and a question every new engineer has to ask.

The data backs this up. A study of Google Chrome's codebase across 39 releases from 2010 to 2015 found that "half of the toggles surviv[ed] 12 or more releases". Even release toggles, the type meant to be short-lived, lingered: "53% still exist after 10 releases indicating that many linger as technical debt" (Rahman et al., MSR 2016). Chrome ships roughly every six weeks, so ten releases is more than a year.

Small teams don't escape this. They just have fewer people who remember why a flag exists.

Step 1: Decide which kinds of flag you're running

Start by naming your flag types, because type decides lifetime, owner and default. Hodgson's four categories are still the clearest split (martinfowler.com, 2017):

TypeWhat it's forExpected lifetimeWho owns it
ReleaseShip incomplete code dark, turn it on when readyDays to weeksThe team shipping the feature
ExperimentA/B or multivariate testsUntil the test reaches significanceProduct or growth
OpsKill switches, load shedding, degrading a costly pathLong-lived by designOn-call / platform
PermissionPlan tiers, beta access, per-customer entitlementsPermanentProduct, often billing

The mistake is treating all four the same. A release toggle that's still around after 90 days is debt. A permission flag that's been around for 90 days is just how your pricing works, and it probably belongs in your entitlements model rather than a flag tool at all.

Unleash bakes this split into the product. Its default expected lifetimes are 40 days for release and experiment flags, 7 days for operational flags (short-lived technical changes such as migrations), and permanent for kill switches and permission flags. Unleash's kill-switch type is closest to what Hodgson calls ops toggles. Past that point, "Unleash marks all flags as potentially stale automatically" (Unleash docs). Whatever tool you use, copy the idea: a flag's type should set its expiry date when it's created.

Step 2: Choose local evaluation for anything on the server

Evaluate flags in your own process against a ruleset cached in memory. Don't call a flag service over the network on every check. This one choice decides your latency, your outage behaviour and your privacy story.

The major platforms agree. LaunchDarkly's server-side SDKs receive "the complete ruleset associated with an SDK key when they initialize" and then evaluate "using their cached ruleset" through "an in-process flag evaluation algorithm", with updates pushed over a persistent connection (LaunchDarkly docs). Unleash's backend SDKs "cache all feature flag data in memory, applying activation strategies locally", and as a side effect "your user data remains within your application and is never shared with the Unleash server" (Unleash docs).

Flagsmith shows both modes, which makes the trade-off clear. In remote mode, the SDK makes a blocking network request every time you fetch an environment's flags. In local mode, the SDK polls for changes "every 60 seconds (by default)" (Flagsmith docs).

Here's how the two models compare:

Local (in-process) evaluationRemote evaluation
Latency per checkMicroseconds, a memory lookupA network round trip
Flag service outageLast known rules keep workingEvery check falls back to defaults
User dataStays in your processSent to the flag service
Change propagationStreaming (seconds) or polling (up to the interval)Immediate
Best forBackend services, workers, APIsBrowsers and mobile, where you can't ship the full ruleset

Browsers and mobile apps are the exception. You can't ship your full targeting rules (including other customers' IDs) to a client device, so client-side SDKs "delegate the flag evaluation to LaunchDarkly on behalf of a specific evaluation context" (LaunchDarkly docs). That's fine. Just know which of your surfaces depends on the network to answer a flag check.

Step 3: Define kill-switch semantics before you need them

Every flag needs an explicit answer to one question: what does this code do when the flag system can't answer? Decide it when you create the flag, not during an incident.

The OpenFeature specification turns this into a contract. Requirement 1.4.10 says evaluation methods "MUST NOT throw exceptions, or otherwise abnormally terminate. Flag evaluation calls must always return the default value in the event of abnormal execution". The default value is also a required argument on every call (OpenFeature spec). Flagsmith has the same idea in a default flag handler, which runs "when a flag cannot be found or if the network request to the API fails" (Flagsmith docs).

The design work is in choosing those defaults:

  • Release flags default to off. If the flag system is down, users see the old, proven behaviour.
  • Kill switches default to "feature enabled". A kill switch exists to turn something off in an emergency. If the flag service is unreachable, you don't want every kill switch to fire at once and take out your recommendation engine, search and billing jobs together.
  • Permission flags default to the most restrictive plan. Failing open on entitlements gives paid features away. Annoying for customers is better than wrong.
  • Ops flags that shed load default to normal operation, for the same reason as kill switches.

Then test it. Point a staging environment at an unreachable flag endpoint and check the app starts, serves traffic and behaves as your defaults say. It's a short exercise, and it catches the cases where an SDK blocks start-up waiting for its first ruleset.

Step 4: Treat every flag change as a production deploy

A flag change can alter production behaviour as much as a deploy can, without a pull request, a review or a CI run. Your audit trail has to fill that gap.

At a minimum, record for every change:

  1. Who changed it (a named person, never a shared account)
  2. What changed, as a before-and-after diff of rules, not just "updated"
  3. When, with timestamps you can line up against your incident timeline
  4. Why, as a required free-text field, ideally with a ticket link

Then send those events where on-call already looks. A flag change that shows up in the same Slack channel or observability timeline as deploys is much easier to spot. "What changed just before this started?" is an early question in any incident, and it belongs in your incident response runbook. If the answer is spread across two tools, you'll find it late.

For production environments, add approvals on high-risk flags (anything touching billing, auth or data deletion). Most hosted tools support this. If you build your own, budget for it explicitly. It's easy to defer.

Step 5: Never reuse a flag

Knight Capital is the clearest example of why this rule exists. Starting on 27 July 2012, Knight rolled out new code for its Retail Liquidity Program (RLP). According to the SEC, "the new RLP code also repurposed a flag that was formerly used to activate the Power Peg code". Power Peg was functionality Knight had stopped using in 2003, but it "remained present and callable" (SEC Release No. 34-70694, 2013).

A technician didn't copy the new code to one of the eight servers, and no second person reviewed the deployment. When trading opened on 1 August, orders carrying the repurposed flag reached that eighth server, which ran the dead Power Peg logic instead. Over a 45-minute period, it "routed millions of orders into the market" and "obtained over 4 million executions in 154 stocks for more than 397 million shares". In the end, "Knight lost over $460 million from these unwanted positions" (SEC, 2013).

Two lessons fall out of that, and both apply to a small team:

  • Flag keys are single-use. Once a flag is retired, its name is retired too. Enforce this in tooling by keeping archived keys and refusing to create a new flag with an old key.
  • Retiring a flag means deleting the code behind it. The flag wasn't the only problem at Knight. Nine-year-old dead code was still callable. Turning a flag off isn't cleanup.

Step 6: Make cleanup part of the definition of done

Flag cleanup only happens when it's attached to something that already happens. As a separate "tech debt sprint", it won't. Hodgson describes the practices that work: some teams add "a toggle removal task onto the team's backlog whenever a Release Toggle is first introduced", others set expiry dates, and some go as far as "'time bombs' which will fail a test (or even refuse to start an application!) if a feature flag is still around after its expiration date" (martinfowler.com, 2017).

A workable version for a small team:

  1. Create the removal ticket with the flag. Same PR, same person, linked from the flag's description.
  2. Set an expiry from the flag type. Release flags get 30 to 40 days by default. Longer needs a written reason.
  3. Fail CI on expired release flags. A soft warning first, then a hard failure after a grace period. A lint rule or a short script that compares flag keys in code against expiry dates in your flag tool will do. If you run a monorepo, make it a cached task alongside lint and tests.
  4. Report stale flags weekly to the owning team, not to a central platform group that can't judge whether a flag is still needed.

Automation helps at scale. Uber's Piranha tool generates code that removes stale flags automatically. Between December 2017 and May 2019 it "generated code cleanup diffs for 1381 flags (17% of total flags)", and 65% of those diffs "landed without any changes" (Ramanathan et al., ICSE-SEIP 2020). Uber's write-up says the tool removed "around two thousand stale feature flags and their related code" (Uber Engineering, 2020). A small team doesn't need Piranha. The lesson is that even a company with dedicated tooling teams found flag cleanup was skipped often enough to be worth automating.

Should you build or buy a feature flag system?

For most small and mid-sized teams, buy, or self-host an open-source tool, and put OpenFeature in front of it. Build only when flags are part of your product itself, for example when your customers configure rollouts inside your platform.

A basic flag store looks like a weekend project: a table, an admin page, an isEnabled() function. The cost is in what comes after:

  • Streaming updates to every process, with reconnect logic
  • Percentage rollouts with stable hashing, so a user doesn't flip between variants on each request
  • Targeting rules, segments and per-environment config
  • Audit trails, approvals and stale-flag reporting
  • SDKs for every language and runtime you use, including mobile

That's a product, and someone has to own it. It's the same trade-off we cover in our build vs buy decision framework for CTOs. Unless flagging sets you apart from competitors, it scores low on strategic importance, and the five-year cost of running a home-grown system usually exceeds the licence fee you were trying to save.

OpenFeature lowers the cost of picking wrong. It's a CNCF project, accepted in June 2022 and incubating since November 2023 (CNCF). Your code calls a vendor-neutral evaluation API. A provider, "an SDK-compliant implementation which resolves flag values from a particular flag management system", plugs in underneath (OpenFeature glossary). Hooks work "similarly to middleware in many web frameworks" (OpenFeature spec), so audit logging, metrics and validation live in one place rather than at every call site. If you outgrow a vendor, or a vendor raises its prices, you swap the provider rather than rewriting every flag check.

Three flag-system patterns worth revising

Three setups worth changing early:

Permission logic living in the flag tool. It's convenient at first. Then billing and the flag tool disagree about what a customer has paid for, and nobody knows which one is right. Keep plan entitlements in the application's own data model. Flags handle rollout of those entitlements, not the entitlements themselves.

The same default for every flag. "Off" looks like the safe choice everywhere until you think through a flag service outage: every kill switch fires at once. Choose defaults per flag type instead: release flags off, kill switches on, entitlements restrictive.

Cleanup treated as housekeeping. Housekeeping rarely makes it into a sprint. Tying the removal ticket to the PR that adds the flag is far more likely to stick.

Frequently asked questions

What is the difference between a feature flag and a config value?

A config value usually changes how the system runs, such as a timeout or a pool size, and lives with the deploy. A feature flag changes what users experience. It can be targeted to specific users or segments, and changes at runtime without a deploy. The line blurs with ops flags, which is why they need the same audit trail as any other flag.

How many feature flags is too many?

There's no universal number. A better measure is the share of release flags past their expiry date, and whether that share is rising. Left alone, flag counts grow: Rahman et al. found that only 20% of Chrome's release toggles had actually been removed during the study period (MSR 2016).

Should feature flags be evaluated on the client or the server?

Evaluate on the server where you can, in process against a cached ruleset. Use client-side evaluation for browsers and mobile apps, where you can't ship your full targeting rules to the device. In that case the vendor evaluates for one user and returns only the results.

Do we need OpenFeature if we only use one vendor?

You don't strictly need it, but it's cheap insurance. It gives you one evaluation API across languages and a standard place for hooks such as audit logging. Changing vendors later becomes a provider swap rather than a codebase-wide rewrite.

The flags you'll regret are the ones nobody owns

Feature flag system design is mostly about ownership. Give every flag a type, an owner, a default and an expiry date. Evaluate in process. Log every change like a deploy. Never reuse a key, and delete the code when you retire a flag. Do those things and the tool you pick matters much less.

If you're designing a flag system, or cleaning up one that's grown past what your team can reason about, talk to us. We help B2B teams set up delivery practices that still work eighteen months later.

Start a conversation

Want to talk about this?

We are happy to discuss the ideas in this note — or where you see things differently.

Feature flag system design: rules you won't regret in eighteen months — DevLume