Aerospace engineering + CS @ Northwestern Polytechnical University · grad-school-bound (2027) · building with AI agents, every day.
I design, run, and debug AI agent systems. Open to remote AI work — agent engineering, AI testing & eval, prompt systems.
01 · agent-workbench
A local-first desktop app for the lifecycle of AI agent projects: idea → building → testing → live pipeline, versioned prompts with line-level diff & rollback, offline-first sync through a ~50-line Cloudflare Worker. Tauri 2 · Rust (three-layer backend + versioned SQLite migrations) · React · ships macOS / Windows / Ubuntu. Dogfooded daily with 12 real projects.
02 · agent-eval-kit
A minimal, offline-reproducible eval harness for LLM apps — a prompt regression gate. 20 Chinese customer-service cases, two prompt versions, rule-based scorers that handle negation (没有优惠券 ≠ violation) and A-not-A questions (能不能用 ≠ 不能用), bad-case reports, 18 unit tests, CI. Verified against DeepSeek: the first real run found three bugs in the harness itself — fixed, each with a regression test.
03 · jungod.cn — a photograph diary A styled-to-the-editorial-monochrome photography site built end-to-end: Cloudflare Pages + Functions + D1 + R2, dual-language, an admin backend, and a photo-day timeline counting from the day I started learning. Photos carry coordinates and show on a world map. Designed, built, deployed, and kept in sync with my Lightroom exports by scripts — by an AI-assisted single developer.
04 · dsh-anchored-standard · ★ 24 A two-phase preset for the DeepSeek Harness agent: minimal-aligned bootstrap, then a full tool suite after the first call. Published as a plugin in the DSH community.
- Ship full apps quickly — two shipped projects above, each under a week of focused work.
- Run agents at scale — ~8B tokens/month on one machine, 160–250B/month across both; I know where real agent workflows break: context bloat, prompt drift, eval gaps.
- Test & evaluate AI — prompt regression, Bad-Case loops, model evaluation (Promptfoo / custom tooling).
- Robotics / SLAM / design — quadrotor autonomy (ROS 2), visual & point SLAM pipelines, and contracted photography with Visual China Group; the aesthetics leak into everything.
Currently: building toward remote AI roles · dogfooding my own agent tooling · extending the eval kit with LLM-as-judge and CI threshold gates.
Reach me via GitHub. Collaboration — and remote AI work — welcome.

