Back to Projects

Magpie

An AI-native finance workspace — a modelling grid built on named variables, agents that propose instead of write, and a reconciliation engine with zero false matches.

Hackathon build — team project
  • Next.js
  • Prisma
  • PostgreSQL
  • LangGraph
  • AWS
Magpie

Every company runs on a spreadsheet that one person understands. A formula points at H47, someone inserts a row, H47 is now a different number, and nobody notices for three weeks. At quarter end the finance team spends more time repairing the model than thinking about the plan. Magpie is built on the bet that the repairing shouldn't be a human job any more.

What it is

Magpie is a single workspace where a company's real data, its forecast, and AI agents that can work on both live together. It replaces three things at once: the planning spreadsheet, the CSV exports copy-pasted between systems, and the junior analyst who spends two days making a chart.

Variables, not cells

The modelling grid at /workspace is made of named variables — "Enterprise ARR", "Churn %" — each with its own formula, unit, trend and history. Formulas reference names rather than coordinates, so Revenue - Costs keeps working no matter how many rows anyone inserts. The H47 bug simply can't happen.

Switching grain between month, quarter and year keeps rollups correct by construction, because every variable knows how it's allowed to aggregate: a currency sums, a percentage doesn't. Scenarios duplicate the plan so what-ifs sit side by side instead of living in model_v4_FINAL_final.xlsx.

Every edit — typed by a person or proposed by an agent — is a command that carries its own inverse. That makes undo and the audit log the same mechanism seen from two sides: you can always answer who changed what, when, and what it was before.

Agents that propose, never write

The agent answers questions like "why did gross margin slip in Q3?" by writing a visible plan, then delegating to two read-only subagents — a model analyst and a data analyst — so a 24-month sweep fills their context windows rather than the one holding the final answer. Reasoning streams back token by token.

It ends with a proposal, not a write. Any write tool halts the LangGraph run and waits for a human to approve before it executes at all. Safety is a state the graph can't leave without a person in it, not a politely worded prompt.

Reconciliation, with hard numbers

The part with a scoreboard is payment reconciliation: proving that every bank credit matches a specific gateway settlement despite mangled references, batched payouts, fees deducted mid-flight and TDS withheld. An LLM comparing thousands of amounts is right almost always — the worst possible property for a system deciding whether two numbers are equal.

So a deterministic matcher does the money maths — reference, amount, date window, fee tolerance, subset-sum for batched deposits — with no model anywhere in the file. Only the ambiguous residue escalates to an adjudication tier, whose structured answer is then re-checked by a plain TypeScript gate that recomputes the arithmetic and rejects anything that doesn't tie.

Against an 11,258-record synthetic batch with a held-out answer key: 100% precision, zero false matches, 98.6% match rate. The remaining 1.4% is left for a human on purpose.

Shipping it

Magpie runs on AWS: the Next.js app and Caddy share an EC2 box, Postgres sits on RDS with no public IP, secrets come from Secrets Manager through a scoped IAM role, uploaded statements are archived to a private versioned S3 bucket, and CloudWatch alarms page through SNS on 5xx spikes. Most of the time went not into the diagram but into the infrastructure details that look like app bugs and aren't.