ANZHE DONG
[01] · Selected work

Production agents,
shipped solo & in flight.

Founder-built and founder-shipped. Five pieces of work that show how I think about agentic systems — from Hale, the messaging-first product I run today, down to the building blocks that hold agents up.

[01]The portfolio

Five projects.
One throughline: ship it, then prove it.

[01]Village Hale Technologies Inc.·
in flight

Village Hale — a number your family texts

My current focus. A messaging-first family chief of staff — parents text one number and Hale takes over the invisible admin: it watches city registration windows across 15 GTA municipalities (the 7 a.m. openings that fill before breakfast), plans the week, and graduates from suggesting to executing as it earns trust, action by action — an independent reviewer agent gates every act, and everything leaves an immutable receipt. No app required; the dashboard is just the receipts room. Privacy is the product — PIPEDA and Quebec Law 25 compliant by default, teen data redacted by construction. Live on a real number since August 2026. Empty repo to production, solo.

Empty repo → prod

solo

Registration radar

15 cities

Compliance

PIPEDA · Law 25

Agent autonomy

earned, gated

  • Agentic systems
  • Privacy-first design
  • Messaging-first consumer AI
Read case study
[02]Settled (TripFix)·
shipped

TripFix — autonomous flight-claim co-pilot

An AI co-pilot for flight-delay refund claims. Reads boarding passes and airline emails, drafts the rebuttal letter, escalates only when uncertain. Built as a small team of specialised agents — not one monolithic prompt — so each piece is testable, swappable, and auditable on its own.

LLMs orchestrated

14+

Eval dimensions

5

Citation grounding

deterministic

Headcount in AI

1

  • Multi-agent systems
  • Eval harnesses
  • Vision reasoning
Read case study
[03]Settled (TripFix)·
shipped

Cursor Cloud Agent v1 — conversation timeline rebuild

A flight recorder for cloud agents. Stitches prompt, thinking, and tool calls into a single replayable timeline — so any agent run is auditable in under a minute. Design call: optimise for the operator first, not the model.

Stream types unified

3

Replay fidelity

100%

PRs merged solo

9

  • Agent observability
  • Tool-use traces
Read case study
[04]Settled (TripFix)·
in flight

Agentic preparation checklist

An agent that reads a case and figures out what’s missing. Instead of one giant ‘knows everything’ prompt, it loads short markdown skills on demand for the stage it’s in. Cheaper inference, sharper answers, knowledge anyone on the team can edit in a text file.

Skills authored

12

Tools wired

7

Snapshot evals

passing

  • Agent design
  • Skills-as-prompts
Read case study
[05]Settled (TripFix)·
shipped

LLM-as-judge evaluation framework

The quality bar for every AI change we ship. Five automatic judges grade each output on truth, sourcing, tone, completeness, and safety. New prompt scores worse than the live one — the deploy is blocked. The only reason daily prompt iteration is safe at production scale.

Evaluators

5

Daily judged samples

hundreds

Regressions caught pre-deploy

many

  • Evals
  • Production safety
Read case study