Arrange Act Assert

Jag Reehals thinking on things, mostly product development

Tag: testing

7 posts tagged “testing”.

Test WebMCP tools like an API (with Playwright)

07 Sep 2026

The Chrome team's best practices and use cases guides end on evaluation tests.

Evals cover the agent's judgement. Did it pick the right tool, did it fill the form the way the user meant, did it stop when it should have asked. You can't hard-code those outcomes, which is why the guides send you to evaluation-driven development.

Underneath that sits a layer that is deterministic and dull. Once your page calls registerTool, it advertises callable functions with typed schemas to whatever agent is driving. That's an API. Get an enum wrong and an agent walks a path your buttons never allow, and the eval that catches it will report that the model got confused. Your UI tests will never see it.

I measured what Chrome 152 does when an agent calls these methods.

The default lane on the command line: the six WebMCP tests, plus the two auth setup tests the chromium project depends on, passing in under three seconds

Read More →

fn(args, deps) Is Not Anti-Mock

23 Mar 2026

People hear "don't use vi.mock" and assume the alternative is "don't mock anything." That is not the argument.

fn(args, deps) is not anti-mock

The real distinction is simpler: mock explicit collaborators, not implicit imports.

Read More →

fn(args, deps) — Bringing Order to Chaos Without Breaking Anything

18 Mar 2026

There’s a major disconnect in AI-assisted development right now. Most of the conversation assumes you’re building something new, or working from the kind of clean, stable foundation that barely exists in real engineering teams.

The reality is that most engineering teams live in legacy systems under high load, with god classes, global singletons, and console.log as observability. The kind of code where every change is a gamble.

This post shows what happens when you apply fn(args, deps) and autotel to those codebases. fn(args, deps) creates the seam for safe change; production telemetry captures the behavioural record that survives when every other spec has decayed.

John Bercow shouting "Order!" in the House of Commons

To prove the point, we’ll do this in plain JavaScript, not TypeScript.

Read More →

If You Only Enforce One Rule for AI Code, Make It fn(args, deps)

04 Mar 2026

AI coding agents produce code faster than you can review and understand it.

One pattern works in both new and legacy codebases because you can adopt it incrementally, without breaking callers.

fn(args, deps) is all you need

For business logic, treat every function as having two inputs: data (args) and capabilities (deps).

Without a clear constraint, generated code becomes harder to reason about: dependencies disappear, side effects spread, composition gets messy. This is why visible structure is essential.

fn(args, deps) is that constraint

For existing code, start with:

functionName(args, deps = defaultDeps)

Dependencies are explicit, not hidden.

There’s no framework and no package to install.

Read More →

Message Isolation? Autotel Makes Tenant Context Flow

17 Jan 2026

The Signadot team is spot on in their Testing Event-Driven Architectures with OpenTelemetry post.

Message isolation using a shared queue: propagate tenant ID in Kafka message headers; consumers use tenant ID for selective message consumption.

They make the case that infrastructure duplication is expensive. Instead of separate Kafka clusters per environment, use tenant ID filtering on a shared queue. Instrument producers and consumers for context propagation.

We've all been there: maintaining four "identical" Kafka setups that slowly drift apart.

Their key insight:

Requires modifying consumers and using OpenTelemetry for context propagation.

But there's still a gap...

Read More →

Request-Level Isolation? Autotel Propagates Context Automatically

16 Jan 2026

The CNCF team is spot on in their Testing Asynchronous Workflows using OpenTelemetry and Istio post.

Request-level isolation is the most cost-effective approach.

They make the case against duplicating infrastructure for testing. Instead of spinning up separate Kafka clusters per tenant, use OpenTelemetry Baggage to propagate tenant ID through async flows. Consumers filter by tenant ID. Istio handles routing.

We've all been there: every team has their own "staging Kafka" and costs balloon.

Their key insight:

Use OpenTelemetry Baggage to propagate tenant ID through sync and async. When publishing to Kafka, producers inject trace context (including baggage) into message headers; consumers extract and make routing decisions.

But there's still a gap...

Read More →

Why Engineers Should Try to Reproduce Production Issues Locally

23 Jan 2023

As engineers, one of our primary responsibilities is to ensure that the systems we build are stable and reliable.

However, despite our best efforts, issues and issues will inevitably arise in production environments. When this happens, it can be tempting to try and patch the problem and move on quickly.

However, recreating production issues locally is a critical step in the debugging and resolution process.

In this post, I'll explain the benefits of reproducing production issues locally.

engineering examining an engine

Photo by Aaron Huber on Unsplash
Read More →