Dave Farley tells a story about an organisation that took months to ship a release. Staged, layered, careful. While someone walked him through the process, he asked what happens when production breaks and a fix has to go out now.
"Oh, we can do a type seven release in under an hour."
So why not run every release as a type seven?
"We couldn't possibly do that. The risk is too high."
Type seven was their name for skipping every check the normal process existed to run, shipping the diff, and hoping. That organisation kept two paths to production: one so slow nobody could learn from it, and one so dangerous they saved it for emergencies.
Your observability has a type seven too.
You know the shape of it. Production is misbehaving, your dashboards answer none of the questions you have, so someone adds a log line and ships it. Waits for the deploy. Reads the output. Adds another log line. Or flips debug logging on across a service for twenty minutes and flips it off before the bill arrives. Or attaches a debugger to a live process and holds their breath.
Nobody files that under "the normal process failed". It goes under incident response, where uncomfortable things go to be forgiven.
The talk this story comes from is worth watching in full:
Organisations that get this right never need a type seven. Releasing is already cheap and fast enough that even a real emergency uses the normal path, with the safety still in it. Fast and safe stop being a trade-off. Two paths exist when you plan your way through a system that only yields to learning.
There’s a major disconnect in AI-assisted development right now. Most of the conversation assumes you’re building something new, or working from the kind of clean, stable foundation that barely exists in real engineering teams.
The reality is that most engineering teams live in legacy systems under high load, with god classes, global singletons, and console.log as observability. The kind of code where every change is a gamble.
This post shows what happens when you apply fn(args, deps) and autotel to those codebases. fn(args, deps) creates the seam for safe change; production telemetry captures the behavioural record that survives when every other spec has decayed.
To prove the point, we’ll do this in plain JavaScript, not TypeScript.
Inject trace context on the producer, extract on the consumer; use PRODUCER and CONSUMER span kinds; set semantic conventions (messaging.system, messaging.destination.name, messaging.operation, Kafka partition/offset/consumer group).
They show the raw OpenTelemetry code. It's comprehensive. It's also verbose. Every team ends up re-implementing the same patterns: inject, extract, span kinds, semantic attributes, error handling.
We've all been there: copying "best practice" code from blog posts and adapting it for our broker.
Their key insight:
For batch processing, use a batch span with links or child spans to contributing traces.
Message isolation using a shared queue: propagate tenant ID in Kafka message headers; consumers use tenant ID for selective message consumption.
They make the case that infrastructure duplication is expensive. Instead of separate Kafka clusters per environment, use tenant ID filtering on a shared queue. Instrument producers and consumers for context propagation.
We've all been there: maintaining four "identical" Kafka setups that slowly drift apart.
Their key insight:
Requires modifying consumers and using OpenTelemetry for context propagation.
Request-level isolation is the most cost-effective approach.
They make the case against duplicating infrastructure for testing. Instead of spinning up separate Kafka clusters per tenant, use OpenTelemetry Baggage to propagate tenant ID through async flows. Consumers filter by tenant ID. Istio handles routing.
We've all been there: every team has their own "staging Kafka" and costs balloon.
Their key insight:
Use OpenTelemetry Baggage to propagate tenant ID through sync and async. When publishing to Kafka, producers inject trace context (including baggage) into message headers; consumers extract and make routing decisions.
Traces break at queues unless you extract context from message headers and put it in the appropriate context.
They walk through the real pain: stateful processing loses trace context in caches, Kafka Connect can only do batch-level tracing, and every team ends up writing custom interceptors and state store wrappers.
We've all been there.
Their key insight:
In Kafka Streams and Kafka Connect this often means manual work: interceptors, state stores, batch spans, or extending tracing logic to extract from headers.