# Five cents

I asked my agent for the price of VXUS, an international stock ETF trading around $87. The first search result it found was `VXUS260828C00089500`, an options contract expiring that afternoon, priced at five cents.

It didn't use that number. But it could have. And if it had, every calculation downstream would have been confidently, fluently wrong, and nothing in the output would have looked unusual.

That near-miss is the reason I built the rest of the project the way I did.

## What I was building

A portfolio rebalancing agent, for the WeMakeDevs Agent Harness Hackathon.

The job is unglamorous. You pick an allocation: 50% US stocks, 25% international, 21% bonds, 4% property. Prices then move at different rates and your actual mix drifts away from the plan. Something has to pull current prices, compute each holding's weight, decide which have moved far enough to matter, and size the trades that would fix it.

Robo-advisors like Wealthfront do this. It's a solved problem, algorithmically. Which made it a good hackathon project for exactly one reason: I already knew what the right answer looked like, so I could tell when the agent got it wrong.

It got it wrong a lot.

## The rule I ended up with

**The model orchestrates. The sandbox computes. The human decides.**

The model reads what you typed, decides which tools to call, and explains what comes back. It does not calculate anything. Every number in the output traces to either a market data server or a tested Python script.

This isn't a stylistic preference. Language models are unreliable at arithmetic in a specific and dangerous way: they fail silently and fluently. A wrong portfolio weight looks exactly like a right one. There is no stack trace, no red text, nothing to catch your eye. Just a number that is plausible and false.

So the model doesn't get to produce numbers. It gets to route them.

## Three things that broke

**"OpenAI-compatible" is a leaky abstraction, and it leaks on agents specifically.**

I set up Gemini 3 Flash-Lite through Google's OpenAI-compatible endpoint, because the free tier gave 15 requests per minute instead of 5. Plain chat worked perfectly. Every single tool call returned a 400.

The cause took a while to find. Gemini 3 returns a `thought_signature` in an `extra_content` field, and that signature must be replayed on the next request when the conversation includes a function call. Standard OpenAI clients, being standards-compliant, drop unknown fields when they rebuild message history. So the follow-up request arrives without it and Google rejects it. Same bug is filed against VS Code Copilot, OpenAI Codex, Open WebUI, and the OpenAI Python SDK itself.

The symptom is distinctive and worth knowing: the tool *runs*, returns its result, and then the model never produces a final answer. Success, then silence.

What makes it interesting is why it breaks agents and not chat. A chat turn sends your message and gets a reply. An agent turn replays the entire assistant history, including the assistant's own function calls, back to the model on every step. Only agent loops touch the field that gets dropped. Compatibility layers are tested on the simple path.

**An agent with two routes to a goal will sometimes take the worse one.**

I had a web search tool connected from early setup, and later added my own market data server with a `get_prices` function. Both could answer "what does VTI cost."

The agent kept choosing web search. Three separate searches, one per ticker, scraping prices out of page text. That is how it ended up looking at that five-cent options contract.

I sharpened the tool description. I wrote "always use this tool for prices, never state a price from memory" into the docstring, because a tool's description is not documentation, it's a prompt injected into the model's context. It helped. It didn't fix it.

What fixed it was turning off the web search tool. Tool selection is probabilistic, and no amount of prompting reliably beats removing the alternative. If a wrong path exists, some percentage of runs take it.

**Writing code that calls tools is cheaper than calling tools.**

TrueForge has a feature called Code Mode: instead of the model emitting native tool calls one at a time, it writes a script that calls tools from inside the sandbox.

I initially read this as a curiosity. It's actually a cost inversion. With native tool calling, four tool calls cost four model round-trips, because the model has to be re-invoked after each result. In Code Mode, the model writes one script that makes all four calls and prints the combined output. One round trip.

On a free tier limited to five requests per minute, that was the difference between runs that completed and runs that died halfway through with a 429.

## What the code review taught me

I ran every pull request through Qodo before merging. It found nine issues. All nine were real, which was humbling, but four of them were the *same* issue wearing different clothes.

I had written type checks everywhere and finiteness checks nowhere. Prices, weights, cash, holdings: each one verified as a number, none verified as a *finite* number. So NaN sailed through every guard I'd written.

The detail that stung: my target weights had to sum to 1.0, and I checked it. But `NaN < 0` is False, and `abs(NaN - 1.0) > tolerance` is also False. A NaN weight passes the negativity check *and* the sum check. Every guard failed open. One helper function using `math.isfinite` closed all four findings.

The sharpest single finding was different in kind. My drift calculator excluded any holding that wasn't in your target allocation. I'd treated "not in your plan" as "not your problem." Qodo pointed out the consequence: a portfolio half-invested in an off-policy asset reports as perfectly on target, no rebalance required. The one holding a rebalance should obviously flag was the one that silently vanished from the maths.

I'd even noticed the edge case. I'd handled it by appending a warning string. A warning next to a wrong number is still a wrong number.

Looking at the findings together, they share a premise I'd got wrong. I built that MCP server as if a developer would call it. The caller is a language model assembling JSON out of a conversation. Malformed input isn't an edge case. It's the expected traffic. Input validation in a tool server isn't defending against malicious users. It's defending against a probabilistic caller being confidently wrong.

## The part I'd keep

The finance has one detail I like more than anything else I built.

The rebalancing rule uses two tolerance bands: act if a holding drifts more than 5 percentage points from target, *or* more than 25% of its own target weight. In one run, VTI drifted 1.15 points and passed. VNQ drifted 1.65 points and failed.

Nearly identical movement, opposite verdicts, because VNQ's target is only 4%, so a quarter of it is one percentage point. A flat 5-point threshold misses that entirely. A small holding can be nearly wiped out while barely moving the headline number.

That rule lives in a `SKILL.md` file in git, not in a prompt. Change the bands and it's a diff, with a review, in the history. That distinction, policy as a versioned artifact rather than a sentence someone retyped, is most of what a harness buys you over a while-loop around an API call.

## What I'd tell myself on day one

The agent framework isn't the hard part. Getting TrueForge running took one command. Writing an MCP server took an afternoon.

The hard part is deciding what the model is allowed to be responsible for, and then actually enforcing it, because the default behaviour of every one of these systems is for the model to helpfully do more than you meant it to. Mine wrote its own drift calculation twice, in place of the tested one sitting right there, and both times the answer looked completely reasonable.

Which is the whole problem. It always does.
