Ep. 014 - Finding Miscompiles For Fun, Not Profit (AI Infrastructure) | Justin Lebar & Jordan Nanos
Summary
- LLM-assisted fuzzer building turned an uncertain weeks-long engineering task into days of work, with reproducible bugs at the end. Justin Lebar had models build fuzzers for NVIDIA’s closed-source PTX compiler and AMD’s LLVM GPU backend, then let agents inspect LLVM’s AMD GPU and x86 code directly. Manual inspection was “basically impossible” and “would fry your brain,” making newly economical infrastructure review the larger opportunity.
- The most severe x86 finding could split an atomic operation into two non-atomic operations, producing the wrong result when contention exposes it. It is dangerous precisely because “most of the time it’ll work fine”; without contention, users might never notice, while “1% of the time” it fails. Lebar found one, maybe two very high-severity x86 bugs and multiple AMD GPU miscompiles, but refused to infer that AMD is buggier from a tool-dependent sample.
- The fuzzer produced about 40 easily reproducible NVIDIA compiler cases, but fuzzing itself hit diminishing returns. It must generate strange yet valid programs, avoid undefined behavior, compare behavior before and after compilation, and stop rediscovering known patterns. As exclusions accumulated, the generator became more complex and its random search less productive; of the 4 or 5 cases Lebar examined, each had no undefined behavior and was clearly a bug.
- The economics improved almost immediately: Opus 4.8 plus Claude Code’s UltraCode cut scan token use to roughly one-tenth. Lebar’s initial read was higher-quality but fewer findings, which he preferred because many earlier flags were discarded anyway; a scan that supported a “$10 grand” budget appeared to cost about $1,000 with UltraCode. He guessed the harness drove more of the gain than 4.8 versus 4.7, while emphasizing, “I really have no idea.”
- The model was useful across the entire remediation loop, not merely discovery. It triaged findings, wrote fixes, and checked Lebar’s edits; he changed model-written fixes about 50% of the time, and the model frequently caught bugs in his own work. The resulting QA flywheel still requires token budgets, human filtering, and maintainers willing to land patches.
- Lebar rejects the unproven claim that an “AI slop cannon” necessarily ships buggier software than the “human slop cannon.” He concedes model code is less beautifully designed and less pleasant to read than work from his best collaborators, but has seen no scientific evidence on comparative bug rates. His practical call: “bugs that you find yourself” beat customer-reported fires, even if organizations reward firefighting more than prevention.
Deep dive
1. GPU compilers may be less exercised
Lebar tested two discovery routes: random-program fuzzing against NVIDIA’s
ptxasand AMD’s LLVM GPU backend, then LLM inspection of LLVM’s AMD GPU, x86, and shared backend code.His premise was exposure: mainstream C++ compilers have processed hundreds of millions or billions of lines, while an LLM uses relatively few GPU kernels compared with all the code behind something like Google Search. Less GPU code might mean less-tested compilers.
The results only partly supported that suspicion. AMD yielded several miscompiles, while x86 produced one or possibly two very severe bugs, but Lebar stressed that finding rates reflect the tools and search methods—not necessarily each backend’s underlying quality.
2. Fuzzing finds bugs until its own complexity takes over
A useful fuzzer must generate unusual but well-formed programs without undefined behavior, then establish equivalence before and after compilation—Lebar ran outputs on GPUs, though an interpreter could instead examine compiler output instruction by instruction.
Returns decayed for three reasons: expanding the valid-program universe became harder, random search reached fewer useful edges, and once roughly 40 bugs had been found, known patterns had to be excluded without suppressing nearby unknown ones. NVIDIA’s fuzzer kept rediscovering the same pattern; teaching the model to avoid it did not work well.
3. Agent harnesses made code inspection economically plausible
In December to January 2025–2026, Lebar repeatedly told Codex, “Good job. Keep going,” after it stopped at one or two bugs. New
/goalharness support sustained the work, while better generated code let him “vibe code the whole thing”—an approach that had failed in January.Direct source inspection was the larger unlock. Even when reviewing more than 30 x86 findings, Lebar said he would not have spotted almost any of them if someone had narrowed the search to the relevant 100 lines.
After publication, Opus 4.8 and UltraCode reduced large LLVM scans to “about a tenth of the tokens.” Bug count fell, but apparent quality rose, and Lebar had already been discarding many earlier findings.
His attribution remained tentative: 4.8 did not appear dramatically more intelligent than 4.7, so he suspected UltraCode’s orchestration mattered more. That was an initial impression, not a controlled comparison.
4. Severity becomes real only through triage and repair
The standout x86 bug transformed a specially constructed atomic operation into two non-atomic ones. It could pass unnoticed without contention, then occasionally return the wrong result—the kind of low-frequency failure that makes a miscompile especially damaging.
NVIDIA’s closed compiler yielded about 40 easily reproducible cases. One clear specimen applied “max, subtract” four times and returned the wrong answer, but without source access Lebar could not determine whether the root cause was broad or ultra-specific.
LLVM maintainers reviewed fixes quickly, and Lebar repaired most high-priority x86 bugs he identified, with the models doing much of the identification. Models wrote the patches, but he altered them about half the time and then used the models to catch mistakes in his revisions.
5. Software-quality claims still outrun the evidence
Jordan Nanos framed the tension as “slop AI” producing buggy software, while suggesting it may be more bloated and perhaps worse per line yet better per feature shipped. Lebar’s pushback: no evidence yet shows AI code is buggier, or whether LLMs find its bugs more or less easily.
The broader call was preventive: spend $10,000—or roughly $1,000 for an UltraCode scan—on codebases that matter, then publish more case studies. Lebar wants the method extended to databases, browsers, other compilers, and perhaps NVIDIA assembly with Mythos or 4.8, while cautioning that “the plural of anecdote is not data.”