AccountWard
Turns a conservator’s bank statements into a balanced draft California court accounting, reconciled to the penny.
Open to Summer 2027: SWE · Quant/Finance · Fintech
Triple major in Finance, Computer Science, and Economics at Santa Clara University. I build where they overlap: multi-model AI, fintech infrastructure, and market simulations.
Seven models vote on each claim (simulated); a split vote goes to a human.
01 · Work
Each carries a status badge, and every number links to its source.
All workTurns a conservator’s bank statements into a balanced draft California court accounting, reconciled to the penny.
A local-first AI workbench: model store, chat, a gated coding agent and a multi-agent Council in one Tauri app.
Agents do the work, other models check it. Built nine ways, then measured, with the negative results reported.
Stress-tested a “can’t-hallucinate” model design: 2 typos cut accuracy 82% → 51%, confidence only 0.79 → 0.61.
An incremental game about running pump-and-dumps, on a deterministic order-book sim that counts every victim.
02 · Research
Calibration under noise, and whether agreement means correctness.
All researchTyped-decision models answer in fixed-shape fields, such as a yes or no, one option from a list, or a score, in a single forward pass, and they are sold as unable to hallucinate. I tested what that guarantee leaves out. BEMA is an open, from-scratch reproduction of the interface: one small transformer encoder with binary, 77-way and regression heads, trained only on human-labeled data and calibrated with temperature, Platt, isotonic and split-conformal methods. On clean held-out text it looks trustworthy: temperature scaling cuts the binary head’s expected calibration error from 0.0159 to 0.0046. Then I added two adjacent-character typos per query, the ordinary fast-typing kind. Accuracy on the 77-way head fell from 81.9% to 51.2% (n = 3,080), while mean confidence fell only from 0.791 to 0.612. The same gap appears on a second, 151-way dataset. Retraining on typo-augmented data halves the accuracy drop, and the gain carries over to typo types it never trained on, but it does not close the gap. The output guarantee is real: the model cannot emit an answer outside its schema. It can still be confidently wrong on input squarely inside its own domain. A noise-robustness number belongs next to ECE whenever a model’s confidence is sold as a safety signal.
Two adjacent-character typos per query cut the 77-way head’s accuracy from 81.9% to 51.2% (n = 3,080), while mean confidence fell only from 0.791 to 0.612.The gap reproduced on CLINC150 with a separate model and tokenizer: accuracy 71.80% → 45.04%, confidence 0.7514 → 0.5615 (n = 5,500).Typo-augmented retraining cut the accuracy drop from 31.4 to 16.0 points at no cost to clean accuracy, and was 11 to 15 points more accurate on typo types it never trained on, but did not close the gap.
Read the paper
A language model is most dangerous when it is confidently wrong, and its own confidence is a weak guard. This position paper describes an architecture I built nine times between April and August 2026. Agents do the work, deterministic checks that no model can vote on test it, a council of models from other labs challenges it, and a human takes every split decision. From those builds come three working theories: adjudication beats synthesis, the model that writes should not check its own work, and agreement is not proof. Only the third has data behind it so far. In a two-model cross-check of 94 form-field mappings, all 3 known errors sat inside the 7 disagreements, though the agreements were never audited. On 20 planted citations, a string matcher and an LLM judge agreed on 19 and were both wrong on 2 of those, so their failures overlapped instead of cancelling. Recent work finds that errors correlate across model providers too. I therefore treat independence as a property to measure, not a feature to buy, and propose an evaluation on three tasks with checkable ground truth: error rates when checkers agree and when they disagree, adjudication against synthesis, and self-review against cross-lab review.
In a two-model cross-check of 94 government-form field mappings, all 3 known errors sat inside the 7 disagreements. The 87 agreements were never audited, so this is a floor, not a rate.A deterministic citation checker and an LLM judge agreed on 19 of 20 planted citations and were both wrong on 2 of those 19: their errors overlapped instead of cancelling.Asked the identical question five times, the same LLM judge gave more than one verdict on 2 of 10 citations, and called a fabricated paraphrase present in 3 of 5 tries.
Read the paper
03 · Thesis
–
An open, MIT-licensed reproduction of a typed-decision model design, built to test the claim that a model with no free text to generate can’t hallucinate. I specified the experiments and directed Claude Code agents through the build and write-up.
Pitched Redl to investors in San Francisco (Jul 2026).
Competed on a team with my collaborator (Fluxxion88). We built Intern, which trains an agent like a new hire: from one hand-made example, an LLM loop writes, runs, scores and repairs a script until it matches, then hands back a script with no model inside.
Built Warrant + Forge with my collaborator (Fluxxion88) over two days, with Warrant forked from Concord’s verifier, ingestion and provider layers: an estate-settlement engine where no fact reaches the ledger without a verbatim quote verified in its source document. It didn’t place, but I opened the AccountWard repo the day after.
06 · Lab
Mapping a market before writing more code, and shelving what doesn’t survive it, counts as a result.
Everything I’ve exploredHow far can coding agents take a Roblox game from a written brief? I ran three agent threads in parallel on free-tier models, one genre each, pushed each with the same finish-it prompt, and routed a stronger model to do a verification pass.
→ Generation is the cheap part: each thread came back as a Rojo-structured Luau codebase with design docs and a built place file. Getting a build to play correctly in Studio is the real work, and so far only BLACKSITE: NULL has been through a Studio playtest.
I bring up ESP32-S3 and LoRa hardware on my own bench. Three ESP32-S3 boards have connected to it, a LilyGO T-Deck among them, on the Arduino and Espressif toolchains with CP210x and CH340K USB-UART drivers, and the Meshtastic 2.7 release is staged for flashing.
→ For bring-up I wrote my own I²C diagnostic in Arduino C++: it pulses the OLED reset line, sweeps addresses 1–126 every 3 seconds and reports each device that ACKs over serial. A working mesh node is the next milestone.
I’ve routed models by task, cost and vendor in three codebases built by AI coding agents under my direction: JobApp tiers work across Claude models, Redl seats each council member on a local or cloud model, and Concord scores models on capability and price.
→ Routing by cost is table stakes. The rule worth keeping is about verification: the model that checks the work comes from a different provider than the model that did it.
After running Concord live, I asked whether it already existed and directed a multi-agent research sweep across six segments of the AI due-diligence market, from PE-diligence software to frontier labs, scoring every Concord feature against them.
→ Multi-model routing, bring-your-own-key desktop and devil’s-advocate review came back commoditised. Two pieces survived: a blocking citation tie-out and per-finding cross-vendor challenge that keeps dissent. Its top recommendation, an adversarial benchmark, is what I built next.
Could M&A due diligence run entirely on open-weight models on the user’s own machine? I forked my Redl app into Quorum and directed Claude Code: five workstream leads, a risk committee with a devil’s advocate, and an IC memo, on Ollama by default.
→ Its gated end-to-end test on a local 7B model produced a PROCEED WITH CONDITIONS memo that flagged a change-of-control clause. Then I forked it again into Concord, to put frontier models from several labs on the same desk.
How much of a 2D action-platformer can a coding agent build in one night from written task specs? I directed OpenAI’s Codex agent through a Godot 4.6 prototype set on a sleeping titan.
→ It got the systems down fast: a tunable movement controller with coyote time, jump buffering and a dash, four enemy archetypes on a shared base class, a boss framework and three zones. No art and no export yet; parked since June 2026.
07 · Contact
Open to Summer 2027 internships in software engineering, quant and finance, and fintech. Open to relocating anywhere and to working on-site.