English · 中文
I build the data systems behind measurable model improvement: failure-driven synthesis, human review with provenance, and regression tests that decide whether a checkpoint is actually better.
My current work sits at the intersection of multimodal learning, post-training, human feedback, and evaluation infrastructure. The common thread is simple: every training example should have a reason to exist, a way to be inspected, and evidence that it helped.
|
Verifiable handwriting synthesis for failure-driven multimodal training and regression.
|
Auditable human-feedback annotation, review, provenance, and frozen data export.
|
A multi-agent adversarial testing harness for defense-policy iteration and regression.
|
- Producing verifiable data for foundation models and post-training systems.
- Multimodal education agents, visual reasoning, and real-task evaluation.
- Human-feedback cleaning, adjudication, attribution, and pre-training quality control.
- Adversarial testing loops that turn model failures into the next regression set.
For research collaboration, private deployment, or multimodal evaluation work, contact me on WeChat: znxzsy. Project-specific questions are also welcome in GitHub Issues.
