Open to Summer 2027: SWE · Quant/Finance · Fintech

Mitansh Mittal

Triple major in Finance, Computer Science, and Economics at Santa Clara University. I build where they overlap: multi-model AI, fintech infrastructure, and market simulations.

Software Engineering Intern, early-stage fintech startup (stealth), Jun–Sep 2026 · SCU ’29

Seven models vote on each claim (simulated); a split vote goes to a human.

FIG 0.1, an illustrative simulation of the Council. Seven models, A to G, vote on a hypothetical claim, “The Q3 accounting balances.” The vote splits 4 to 3, a disagreement score of 0.43, so the claim escalates to a human.
  1. A agrees
  2. B dissents
  3. C agrees
  4. D dissents
  5. E agrees
  6. F dissents
  7. G agrees
  • Triple major: Finance × Computer Science × Economics
  • 2 awards at CruzHacks 2026, UC Santa Cruz
  • NEXUS · solo at SCSP AI+ Expo Hackathon
  • BEMA · 81.9% → 51.2% under 2 typos
  • 997 tests · EXIT LIQUIDITY's order book
  • Warrant · 41 / 41 rewritten quotes caught
  • Pitched Redl to investors in SF (Jul 2026)

In numbers

majors: Finance, Computer Science and Economics
3
projects, each badged by status
44
hackathons entered
6
papers, essays and academic write-ups
8
languages in pushed code, by bytes (mit37, reviewed repos)
6
Sources
  1. mitansh-confirmed 2026-09-22
  2. 44 entries in src/content/work at build
  3. 6 entries with kind: hackathon in src/content/experience
  4. 8 entries in src/content/research at build
  5. src/data/github.json, refreshed at build (scripts/github-stats.mjs)

01 · Work

Five builds across AI, fintech and markets.

Each carries a status badge, and every number links to its source.

All work

Fintech & RegTechin-build

AccountWard

Turns a conservator’s bank statements into a balanced draft California court accounting, reconciled to the penny.

AI Systemsbuilt

Redl

A local-first AI workbench: model store, chat, a gated coding agent and a multi-agent Council in one Tauri app.

AI Systemsresearch

The Council

Agents do the work, other models check it. Built nine ways, then measured, with the negative results reported.

Researchresearch

BEMA

Stress-tested a “can’t-hallucinate” model design: 2 typos cut accuracy 82% → 51%, confidence only 0.79 → 0.61.

Games & Simulationin-build

EXIT LIQUIDITY

An incremental game about running pump-and-dumps, on a deterministic order-book sim that counts every victim.

02 · Research

Two papers, measured on real data.

Calibration under noise, and whether agreement means correctness.

All research

working draftresearch

Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noise

Typed-decision models answer in fixed-shape fields, such as a yes or no, one option from a list, or a score, in a single forward pass, and they are sold as unable to hallucinate. I tested what that guarantee leaves out. BEMA is an open, from-scratch reproduction of the interface: one small transformer encoder with binary, 77-way and regression heads, trained only on human-labeled data and calibrated with temperature, Platt, isotonic and split-conformal methods. On clean held-out text it looks trustworthy: temperature scaling cuts the binary head’s expected calibration error from 0.0159 to 0.0046. Then I added two adjacent-character typos per query, the ordinary fast-typing kind. Accuracy on the 77-way head fell from 81.9% to 51.2% (n = 3,080), while mean confidence fell only from 0.791 to 0.612. The same gap appears on a second, 151-way dataset. Retraining on typo-augmented data halves the accuracy drop, and the gain carries over to typo types it never trained on, but it does not close the gap. The output guarantee is real: the model cannot emit an answer outside its schema. It can still be confidently wrong on input squarely inside its own domain. A noise-robustness number belongs next to ECE whenever a model’s confidence is sold as a safety signal.

Two adjacent-character typos per query cut the 77-way head’s accuracy from 81.9% to 51.2% (n = 3,080), while mean confidence fell only from 0.791 to 0.612.The gap reproduced on CLINC150 with a separate model and tokenizer: accuracy 71.80% → 45.04%, confidence 0.7514 → 0.5615 (n = 5,500).Typo-augmented retraining cut the accuracy drop from 31.4 to 16.0 points at no cost to clean accuracy, and was 11 to 15 points more accurate on typo types it never trained on, but did not close the gap.

Read the paper

position paperresearch

Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systems

A language model is most dangerous when it is confidently wrong, and its own confidence is a weak guard. This position paper describes an architecture I built nine times between April and August 2026. Agents do the work, deterministic checks that no model can vote on test it, a council of models from other labs challenges it, and a human takes every split decision. From those builds come three working theories: adjudication beats synthesis, the model that writes should not check its own work, and agreement is not proof. Only the third has data behind it so far. In a two-model cross-check of 94 form-field mappings, all 3 known errors sat inside the 7 disagreements, though the agreements were never audited. On 20 planted citations, a string matcher and an LLM judge agreed on 19 and were both wrong on 2 of those, so their failures overlapped instead of cancelling. Recent work finds that errors correlate across model providers too. I therefore treat independence as a property to measure, not a feature to buy, and propose an evaluation on three tasks with checkable ground truth: error rates when checkers agree and when they disagree, adjudication against synthesis, and self-review against cross-lab review.

In a two-model cross-check of 94 government-form field mappings, all 3 known errors sat inside the 7 disagreements. The 87 agreements were never audited, so this is a floor, not a rate.A deterministic citation checker and an LLM judge agreed on 19 of 20 planted citations and were both wrong on 2 of those 19: their errors overlapped instead of cancelling.Asked the identical question five times, the same LLM judge gave more than one verdict on 2 of 10 citations, and called a fabricated paraphrase present in 3 of 5 tries.

Read the paper

03 · Thesis

Most finance people can’t ship software. Most engineers don’t understand where the money moves. I’m studying all three sides: finance, computer science and economics.

Finance

  • AccountWard: Turns a conservator’s bank statements into a balanced draft California court accounting, reconciled to the penny.
  • EXIT LIQUIDITY: An incremental game about running pump-and-dumps, on a deterministic order-book sim that counts every victim.
  • BEMA: Stress-tested a “can’t-hallucinate” model design: 2 typos cut accuracy 82% → 51%, confidence only 0.79 → 0.61.

Computer Science

  • Redl: A local-first AI workbench: model store, chat, a gated coding agent and a multi-agent Council in one Tauri app.
  • The Council: Agents do the work, other models check it. Built nine ways, then measured, with the negative results reported.
  • NEXUS: Solo-built wargame adjudicator: five Claude agents debate while a 10,000-run Monte Carlo cascade does the math.

Economics

  • ECON 3: Closing speech for the CON side of a team debate on whether U.S. farm subsidies harm the world’s poorest, reframed around who bears the cost.
  • AI due-diligence market map: After running Concord live, I asked whether it already existed and directed a multi-agent research sweep across six segments of the AI due-diligence market, from PE-diligence software to frontier labs, scoring every Concord feature against them.
  • EXIT LIQUIDITY: Order-book and market-microstructure simulation: An incremental game about running pump-and-dumps, on a deterministic order-book sim that counts every victim.

The overlap

  • The Council: Agents do the work, other models check it. Built nine ways, then measured, with the negative results reported.
  • BlueCollarPal: Procurement and virtual-card control plane for trade contractors: text a request, approve it, pay with a one-time card.
  • Concord: Multi-vendor M&A diligence engine: 11 agent desks, cross-lab challengers and a model-free tie-out on every quote. Shelved after mapping the market.

04 · Experience

Startup timelines and hackathon clocks.

  1. –

    Independent researchBEMA: stress-testing a “can’t-hallucinate” decision model

    An open, MIT-licensed reproduction of a typed-decision model design, built to test the claim that a model with no free text to generate can’t hallucinate. I specified the experiments and directed Claude Code agents through the build and write-up.

  2. Pitched Redl to investorsRedl, my local-first AI workbench

    Pitched Redl to investors in San Francisco (Jul 2026).

  3. Intern, team build with Fluxxion88Loop Engineering Hackathon, AWS Builder Loft

    Competed on a team with my collaborator (Fluxxion88). We built Intern, which trains an agent like a new hire: from one hand-made example, an LLM loop writes, runs, scores and repairs a script until it matches, then hands back a script with no model inside.

  4. Warrant + Forge, two-person team with Fluxxion88Estate-settlement AI hackathon

    Built Warrant + Forge with my collaborator (Fluxxion88) over two days, with Warrant forked from Concord’s verifier, ingestion and provider layers: an estate-settlement engine where no fact reaches the ledger without a verbatim quote verified in its source document. It didn’t place, but I opened the AccountWard repo the day after.

05 · GitHub

7 contributions in the last year.

By bytes pushed

  • Python 52%
  • JavaScript 38%
  • CSS 9%
  • PowerShell 1%
  • Batchfile 1%
  • HTML 0%
  • BEMA

    BEMA: stress-testing a typed-decision 'can't-hallucinate' model: calibration and robustness under realistic input noise

    Python · pushed

  • aerobytes

    SlugBites

    JavaScript · pushed

06 · Lab

Ideas I investigated, and what I concluded.

Mapping a market before writing more code, and shelving what doesn’t survive it, counts as a result.

Everything I’ve explored
  • 2026-09explored

    AI-assisted game production pipeline

    How far can coding agents take a Roblox game from a written brief? I ran three agent threads in parallel on free-tier models, one genre each, pushed each with the same finish-it prompt, and routed a stronger model to do a verification pass.

    → Generation is the cheap part: each thread came back as a Rojo-structured Luau codebase with design docs and a built place file. Getting a build to play correctly in Studio is the real work, and so far only BLACKSITE: NULL has been through a Studio playtest.

  • 2026-09explored

    ESP32-S3 / LoRa bench

    I bring up ESP32-S3 and LoRa hardware on my own bench. Three ESP32-S3 boards have connected to it, a LilyGO T-Deck among them, on the Arduino and Espressif toolchains with CP210x and CH340K USB-UART drivers, and the Meshtastic 2.7 release is staged for flashing.

    → For bring-up I wrote my own I²C diagnostic in Arduino C++: it pulses the OLED reset line, sweeps addresses 1–126 every 3 seconds and reports each device that ACKs over serial. A working mesh node is the next milestone.

  • 2026-09explored

    Model routing in practice

    I’ve routed models by task, cost and vendor in three codebases built by AI coding agents under my direction: JobApp tiers work across Claude models, Redl seats each council member on a local or cloud model, and Concord scores models on capability and price.

    → Routing by cost is table stakes. The rule worth keeping is about verification: the model that checks the work comes from a different provider than the model that did it.

  • 2026-07explored

    AI due-diligence market map

    After running Concord live, I asked whether it already existed and directed a multi-agent research sweep across six segments of the AI due-diligence market, from PE-diligence software to frontier labs, scoring every Concord feature against them.

    → Multi-model routing, bring-your-own-key desktop and devil’s-advocate review came back commoditised. Two pieces survived: a blocking citation tie-out and per-finding cross-vendor challenge that keeps dissent. Its top recommendation, an adversarial benchmark, is what I built next.

  • 2026-07archived

    Quorum

    Could M&A due diligence run entirely on open-weight models on the user’s own machine? I forked my Redl app into Quorum and directed Claude Code: five workstream leads, a risk committee with a devil’s advocate, and an IC memo, on Ollama by default.

    → Its gated end-to-end test on a local 7B model produced a PROCEED WITH CONDITIONS memo that flagged a change-of-control clause. Then I forked it again into Concord, to put frontier models from several labs on the same desk.

  • 2026-06explored

    Colossus Wake

    How much of a 2D action-platformer can a coding agent build in one night from written task specs? I directed OpenAI’s Codex agent through a Godot 4.6 prototype set on a sleeping titan.

    → It got the systems down fast: a tunable movement controller with coyote time, jump buffering and a dash, four enemy archetypes on a shared base class, a boss framework and three zones. No art and no export yet; parked since June 2026.

07 · Contact

Let’s talk.

Open to Summer 2027 internships in software engineering, quant and finance, and fintech. Open to relocating anywhere and to working on-site.

m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h