Recorded run · ubusd · planners on one briefing
Empty next_argv. Four runners still used r2b verify.
Grok 4.6, Codex, Kimi K3, and GLM got the same 24.10.3 ubusd briefing. next_argv was empty. Rules review did not promote strcpy. All four ran r2b verify --import strcpy and named the three dynamic callers. None shelled out to raw r2. Kimi and GLM had no agent CLI here; they got one host-owned tool that can only exec uv run r2b.
The productive loop is not auto-queued decompile. It is: read the memory capsule, run the r2b verifier, then decompile a function VA from that JSON. Empty next_argv is valid. Sidestepping is r2 axt or Ghidra on entry.
01 · SAME CARD
Network ranked first. strcpy sat in a lead capsule.
imports:network 90 lead
entry:entry 89 0x1e18 lead
imports:memory 84 memcpy / sprintf / strcpy lead
next_argv []
Rules review with thesis “put strcpy callers ahead of entry” kept that order. The planner has to notice dangerous_imports.
02 · FOUR RUNNERS
All four stayed on r2b. All four hit the gold sites.
verify strcpy callers then decompile sidestep
Grok 4.6 yes 0x4b50 0x4b6c 0x4ba8 0x4b50 call site no
Codex gpt-5.4 yes same 0x4ba8 call site no
Kimi K3 yes same 0x33c4 r2 fn no
GLM glm-5.3 yes same 0x33c4 r2 fn no
Gold from the Kali ubusd case: those three addresses, third one in fcn.00004a3c. All four named them from r2b verify JSON. GLM answered on the Z.ai coding-plan host, not China paas.
03 · THE HOP
The recorded decompiles used call sites. The CLI now emits function_addr.
Grok and Codex passed a call VA. Kimi and GLM passed r2’s name 0x33c4. Ghidra said no function at either. That was a handoff gap, not a model miss. Verify now includes function_addr. A later live r2b decompile … 0x4ba8 returns Ghidra FUN_00104a3c at 0x00104a3c. Do not score the recorded decompiles as a bake-off.
THE HONEST CLAIM
r2b is useful here because empty next_argv did not send them to r2.
This is not a four-way quality ranking. It is a demo that a planner can keep r2b in the loop on the firmware path that actually works. Keys live in gitignored .env. review --mode llm still cannot run verify.