Why this problem is harder than it looks
MATLAB and C++ occupy adjacent niches in numerical software. MATLAB is where algorithms are prototyped, tuned, and validated against reference data. C++ is preferred in production because it runs far faster. Moving code from the first environment to the second is a routine requirement, and it is consistently underestimated. The syntactic gap between MATLAB and C++ is small enough that a mechanical translation looks feasible, but the semantic distance is larger, and it is concentrated in places that do not show up in a casual reading of the source:
- Indexing and storage order: MATLAB arrays are 1-indexed and column-major. C++ numerical libraries typically are not. A loop that is correct under one convention will read the wrong memory under the other, without any compiler diagnostic.
- Broadcasting and shape promotion: MATLAB expands singleton dimensions implicitly. C++ requires explicit materialization.
- Aliasing and copy semantics: MATLAB arrays behave as if they are copied on assignment. C++ containers do not. A direct translation can introduce aliasing bugs that only appear when an input is reused downstream.
- Numerical library equivalence: MATLAB’s built-in functions are not always bit-equivalent to their nominal C++ counterparts. mldivide, eig, and the various fft variants dispatch to different underlying routines depending on input properties, and a faithful port must either replicate the dispatch logic or accept a tolerance gap.
Correctness is judged numerically, typically with relative error in the range 10⁻⁹ to 10⁻¹² for double-precision pipelines. Errors at that scale are invisible to compilation and to functional tests; they surface only in a regression harness comparing MATLAB-generated reference outputs, and usually several call sites downstream from the actual mistake.
What automated tools currently do
MATLAB Coder is the obvious reference point. It handles the mechanical part of the translation reliably and produces code that is numerically faithful to the source. It is also difficult to use. It carries its own runtime, its own container types, and its own conventions for memory ownership, and it does not attempt to integrate with an existing C++ codebase. For a greenfield deployment, this is acceptable. For incremental migration into an existing codebase, the generated code has to be substantially rewritten before it can be used.
The work that remains after Coder — or after any similarly mechanical tool — is the work that requires judgment: choosing idiomatic C++ constructs, deciding when to deviate from the structure of the MATLAB source, fitting the result into the existing data structures and algebra layers of the surrounding codebase, and identifying the places where the two languages differ in ways that affect correctness. This is the part we set out to automate.
A single agent is not enough
The obvious first attempt is one agent with access to the source, the build system, and the regression harness. It works on small inputs and fails in three predictable ways as inputs grow:
- Context pressure causes silent constraint loss: As the working context fills, instructions given early in the conversation stop influencing the agent’s behaviour. This is difficult to detect from the outside, because the agent continues to produce confident output.
- Self-verification is unreliable: When the same agent writes the code and interprets the test results, it has a degree of freedom it should not. We observed cases where, faced with a persistent tolerance failure, the agent adjusted the test fixture or the tolerance specification rather than the implementation. The behaviour is not adversarial; it is the path of least resistance given the agent’s objective.
- Procedural work consumes the context budget: A non-trivial fraction of the agent’s tokens were spent rederiving repeatable procedures: how to invoke the build, how to parse the test output, how to format a numerical diff. This is deterministic work that does not belong in a reasoning loop.
The workflow we settled on
The architecture we settled on (see diagram) has two commitments: split responsibilities across agents with non-overlapping scope, and move all deterministic work into scripts those agents invoke. The pipeline runs left-to-right from Mapper through to Reviewer, with the Orchestrator coordinating from above:
- Orchestrator — Maintains the migration plan from the source call graph, dispatches work, and tracks completion. Does not write code.
- Mapper — Given a MATLAB construct, returns a structured record of the candidate C++ equivalent and the assumptions under which it holds. Its record of previously found matches feeds the Translator.
- Translator — Consumes Mapper records and applies the Code Rules to produce annotated translation decisions.
- Implementor — Runs scripts for build-system integration and other deterministic or simple procedures.
- Tester — Runs the test-creation and execution scripts, then emits a structured diff against the Matlab reference. Has no capacity to modify the implementation or tolerance spec.
- Reviewer — Reads the final output and produces a report of hardcoded constants, assumed input ranges, and integration concerns. Its output is for human readers only.
The scripts handle reference capture from MATLAB, build invocation, test execution, numerical diffing, and report formatting. They are generated once at the start of a migration and treated as fixed infrastructure thereafter.
The separation between translator and verifier is the part that matters most. It removes the degree of freedom that allowed the single agent to resolve tolerance failures by adjusting the test rather than the code. The verifier’s output is structured, the translator’s input includes that output verbatim, and neither can suppress it.
Where this helps, and where it does not
Constraint drop-out is reduced, because each agent works over a smaller context and a narrower task. It is not eliminated. Constraints that span agent boundaries — such as a naming convention that the translator must apply and the reviewer must check — have to be carried by the Orchestrator, and the propagation is occasionally lossy. Future work may replace the Orchestrator with a script that makes CLI calls to the sub-agents, to minimise those losses further.
Procedural rederivation is eliminated by construction, because the relevant work is no longer done by the agents.
Costs
Total token consumption is higher than the single-agent baseline, by a factor of roughly 3 to 5 in our measurements. The increase is dominated by inter-agent handoffs, each of which carries enough structured context to be useful to the receiving agent. Per-iteration cost is therefore higher, but the number of iterations required to reach a passing state is lower, and the overall wall-clock time is substantially shorter.
Repeatability is partial. Two runs over the same input produce implementations that satisfy the same tolerance bounds but differ in structure and naming. This is tolerable for one-shot migration. It is awkward when a module has to be re-migrated after an upstream change, because the diff is not meaningful at the line level.
Human review remains necessary at the integration boundary. The Reviewer narrows what must be read carefully, but it does not remove the requirement for the parts where judgment was applied. We have not yet seen a case where we would deploy the output without a human reading it.
What we take from this
The single-agent baseline failed in ways that are intrinsic to giving one model the implementation task and the verification task at once, while also asking it to handle the procedural overhead of running the loop. Splitting these into separate agents with non-overlapping scope, and moving the procedural overhead into scripts, is what changed the workflow from an experiment into something we run on real modules. The specific division of labour described above is probably not the only one that works, but the two structural commitments behind it seem to us to be the load-bearing parts.