CI/CD That Fails Fast on Purpose

More from the blog
October 1, 2026

A pipeline that takes forty minutes to tell you someone left a typo in a variable name is not a safety net. It's a queue with extra steps.

Failing fast isn't about being harsh on developers. It's about ordering your checks so the cheapest and most likely failure is caught first, and the expensive ones only run when they've earned the compute.

Ordering is the whole design

Most pipelines inherit their stage order from whatever template the vendor shipped: lint, build, test, deploy. That order is rarely optimised for feedback, and feedback latency is the only metric developers actually feel.

Sort your stages by probability of failure multiplied by cost of failure, not by tradition. The result is almost always the same five rungs.

  • Formatting and lint — seconds. Catches typos, unused imports, and style drift. Costs nothing, fails often. Run it first, and run it on the diff rather than the whole repo.
  • Build or compile — one to three minutes. A dependency that doesn't resolve or a type error should never reach a test runner.
  • Unit tests — a few minutes. Fast, isolated, parallelisable. If your unit suite takes fifteen minutes, it isn't a unit suite.
  • Integration tests — ten to thirty minutes. Real databases, real queues, real network. Expensive by nature, so they should only run on code that already compiles and passes units.
  • End-to-end and deploy previews — thirty minutes and up. The most expensive and the flakiest. Run them on merge, not on every push, and never block a one-line docs change on them.

That ordering alone typically cuts median feedback from tens of minutes to under three, without changing a single test.

Fail fast means fail loudly

Ordering is useless if the runner keeps going after something breaks. Two habits cause most of the damage: steps configured to continue on error, and retry logic that hides real failures behind three green attempts.

Stop on the first failure

Every quality gate should be a hard gate. If lint fails, the job ends there — don't burn twenty minutes of runner time to discover unit tests also fail. With parallel jobs, that means marking downstream stages as dependent and letting the whole run cancel when an upstream one fails.

Kill flaky tests, don't rerun them

A test that fails one run in five is not a slow test, it's an untrustworthy signal. Rerunning it trains the team to press the button again instead of reading the output. Track flake rate, quarantine anything above a couple of per cent into a separate, non-blocking job, and treat a repaired flake as a shipped fix.

Put a clock on everything

Every job gets a timeout. A pipeline with no timeout will eventually hang on a network call and hold a runner for six hours. A build that takes four times its usual duration is a failure condition even when it passes — it's usually a cache miss storm or a test that quietly grew an integration dependency.

A pipeline's job is to be wrong quickly, so that humans have time to be right.

What good looks like in practice

The target numbers are unglamorous. Under three minutes for the first meaningful signal on a pull request. Green or red, never amber. A failing run that tells you which file and which line, with the artefact needed to reproduce it locally. And a main branch that is always deployable, because nothing merges until the fast checks have passed.

Speed comes from structure more than hardware. Caching dependencies, splitting the unit suite across runners, running lint only on changed files, and letting slow tests run on merge all cost less than upgrading the runner tier.

Where Weeltec comes in

We rebuild CI/CD pipelines for teams whose builds have quietly become the bottleneck. That usually means reordering stages, splitting the monolith pipeline into a fast gate and a slow verification path, removing continue-on-error flags that were hiding real failures, and making every job reproducible on a developer's machine. The tooling rarely changes. The order and the discipline do.

The rule

Measure the time from push to first red or green signal, and treat that number as a product metric. If it's longer than a coffee, your pipeline is telling developers to context-switch — and they will, permanently, into something else.

Weeltec builds and tunes CI/CD pipelines for teams that measure their feedback loop in minutes, not meetings. If your builds have become the slowest part of shipping, get a quote and we'll reorder the whole thing.