Every firmware engineer using AI right now has had this moment.
The code compiles. The unit tests pass. You flash it. And the board does something you did not ask for.
It feels like a personal failure. It is not. It is structural, and three weeks ago the structure acquired a name.
The loop that closed
Agentic development transformed how software gets built, and the reason is not that AI writes brilliant code. It is that AI can see what the code does.
An agent writes a function, deploys it to the host, runs it, reads the result, and fixes what broke. All of that happens on a machine the agent can inspect. The feedback is immediate. The loop closes on its own, thousands of times a day, with no human in the middle. That visibility is the whole mechanism. Take it away and you have a very fast typist with no idea whether anything works.
The loop that did not
Embedded teams use the same AI, and it writes firmware well. Register setup, peripheral initialisation, DMA descriptors, the boilerplate that used to eat afternoons. Ask most firmware engineers and they will tell you writing got noticeably faster this year.
Then the code has to run on real hardware. And there, the AI is blind.
It cannot see a signal. It cannot see the timing. It cannot see the 3.3 V rail dip when three GPIOs switch at once, or an interrupt priority that is correct on its own and wrong under load. It cannot put a probe on anything. So a person still flashes the board, still scopes it, still decodes the protocol, still works out what happened. The fast part got faster. The slow part stayed exactly as slow.
That is the broken loop. It is where embedded projects lose their months, and until recently it did not have a name in the wider AI conversation.
What Anthropic did on 27 August
On 27 August, Anthropic opened a research preview of something called the Model Hardware Standard. MHS is a shared specification that lets AI agents discover, understand and safely operate physical devices. If MCP is how an agent talks to software tools, MHS is the same idea pointed at machines.
The first cohort is scientific labs and advanced manufacturers. The examples in the announcement are microscopes, liquid handlers, robotic arms and a laser on a quantum computer. The standard was developed with HHMI Janelia Research Campus. Amazon Web Services, Automata and Danaher are building support. The pitch is blunt: integrating a new instrument into an automated setup used to take weeks or months of bespoke driver work, and Anthropic says MHS reduces it to hours or minutes.
Two things about MHS matter here.
First, it is built on MCP. From the agent’s point of view, an MHS device is not so different from an MCP server. It advertises capabilities, accepts structured calls and returns structured observations. Any team already running agents against MCP surfaces is looking at another MCP surface, not a new stack.
Second, it is a research preview, not a shipped standard. Anthropic plans to open-source the full specification. That has not happened yet. Anyone claiming MHS support today is claiming something the standard’s own authors have not released.
Why hardware-in-the-loop was the first thing people thought of
The most-shared engineering take on MHS came from Dr. Dirk Alexander Molitor, who wrote on LinkedIn that it could become the bridge between engineering AI and physical AI. His first example was hardware-in-the-loop testing: an agent configures a HIL setup, flashes an ECU, executes test scenarios, monitors signals, analyses failures and adapts the next test sequence.
That is the loop. Design, simulation, physical test, learn, redesign. The faster observations from real hardware flow back into engineering decisions, the shorter the development cycle.
It is also, I think, the hardest case MHS will meet, and nobody has written about why.
A microscope reports its state. A liquid handler either dispensed or it did not. These are instruments that expose their behaviour through an interface, and MHS standardises that interface. Embedded firmware under test is different in kind. The device is the thing being changed. Its misbehaviour lives in twelve microseconds of interrupt latency, in a rail that sags for a few hundred nanoseconds, in a protocol edge case the datasheet never mentioned. None of that is exposed as a clean API. The bench has to observe it from outside, with a scope and a logic analyser, on a timeline the firmware itself cannot see.
So MHS does not close the embedded loop by existing. It makes the instruments on the bench legible to an agent. Someone still has to build the loop that uses them, decide what to measure, and stand behind what the measurements mean.
The caveat Anthropic put in its own announcement
Coverage of the MHS preview carried a line that deserves more attention than it got. Early tests showed agents could optimize workflows and generate scripts that ran without a language model in the loop. But the system still struggled with understanding physical cause and effect, so human oversight remains a must.
That is not a limitation to engineer around. It is the design principle.
An agent operating on real hardware has to be watchable, and it has to be stoppable. The agent proposes. The compiler, the probe and the scope decide. The engineer decides whether to trust any of it, and can change it before it runs. Anthropic’s own testing landed on the same conclusion from the other direction: the AI can do the work, and a person still has to be in the loop for the physical part. For embedded, where a wrong write can brick a board, that is not optional.
Where the months actually go
Ask any firmware lead where a project lost its time and the answer is rarely the code.
It is the bench. Wiring a test rig by hand. Reproducing a fault that only appears under load. Decoding a protocol from a raw capture. Waiting on the one engineer who knows how the rig is cabled. Then doing it all again when the pinout changes.
Software teams solved this years ago, because their code runs where every test can be automated and every result observed. Embedded teams never got that, because their code runs on hardware that has to be measured by a person with instruments. AI made the first half of the job fast. It has not touched the second half. That is the gap MHS points at, and it is the gap that decides whether AI in embedded development is a productivity story or a sales pitch.
The question
So here is what we have been asking ourselves at Better Devices, and what I want to put to anyone reading this who builds firmware for a living.
Is anyone actually closing this loop on embedded hardware? Not faster code generation. Not better simulation. The specific problem of an AI-written binary meeting real silicon, with something other than a tired engineer and a scope checking what happened.
Anthropic has named the category. The instruments are about to become legible. The hard part, the part between the generated code and the physical result, is still open.
I have a view on it. I will share it at the end of this week. But I would like to hear yours first, particularly if your team has tried this and it did not work.
Frequently asked questions
What is the Model Hardware Standard?
A specification Anthropic released as a research preview on 27 August 2026 that lets AI agents discover, understand and safely operate physical devices such as microscopes, robotic arms and manufacturing equipment. It is built on the Model Context Protocol and is planned for open-source release.
How is MHS different from MCP?
MCP standardises how an AI agent talks to software tools and data sources. MHS extends the same pattern to physical hardware, standardising how a device exposes its capabilities, state, parameters and safety constraints. An MHS device looks to the agent like another MCP surface.
Does MHS apply to embedded firmware testing?
Not yet directly. The research preview targets labs and advanced manufacturing. Hardware-in-the-loop testing is an obvious application, but embedded firmware is a harder case than a lab instrument, because the device under test is the thing being changed and its faults appear in timing and signal behaviour that no API exposes.
What does AI in hardware-in-the-loop testing actually mean?
An AI agent driving a physical test bench: flashing firmware to a device, applying stimulus, capturing the electrical response with instruments, and checking the result against a specification, with an engineer reviewing and approving what runs.
Why can AI not see what firmware does on the board?
Because the firmware runs on a separate physical device, and its behaviour, signal levels, timing, power draw, is only observable with external instruments. Unlike software on a host machine, there is no built-in feedback path from the running code back to the agent that wrote it.
