AI Researcher & Engineer at BAAI · Previously ByteDance and Meituan
I build dependable foundation models for specialized domains through post-training, retrieval agents, and multimodal reasoning. Explore the selected projects below for papers, code, datasets, and models.
Website · Google Scholar · Hugging Face · Email
- MechVQA / MechVL
ICML 2026— A benchmark and domain-specialized models for understanding mechanical drawings. Code · Data & models - IAR
2026 preprint— Staged post-training to internalize document knowledge, align question answering, and recover general capabilities without retrieval. - SPAR / SPARBench
2025 preprint— Multi-agent scholarly retrieval with an evaluation dataset. Code · Data - SciSage / SurveyScope
2025 preprint— Multi-agent scientific survey generation and evaluation. Code · Data
More work on post-training: Wnuan · RAFT · SFTKey · MoSLD (COLING 2025).
Multimodal retrieval and domain models: ChartWalker · CareBot (AAAI 2025) · Aquila-Med.
Full research index · Research collaboration · Published patent applications
- CCI3.0-HQ
— high-quality Chinese pre-training data · Paper
- IndustryCorpus2
— multilingual multi-industry pre-training corpus
- IndustryCorpus
— the original multilingual multi-industry corpus
- Industry Instruction
— multilingual multi-industry instruction data
These are contributions to projects maintained by their respective organizations and authors:
- FlagAI — Added the Aquila-SQL training, inference, and evaluation example. Contribution
- patent-disclosure-skill — Added inventor-based patent portfolio retrieval from CNIPA's publication system, with parsing, tests, and documentation. Contribution
- paper-reading-skill — Agent skill for visual, source-grounded paper reading: one standalone HTML atlas of methods, equations, and experiments.
- prompt-thinking-toolkit — A thinking-mode router for agents: 12 bounded prompt workflows (Socratic diagnosis, first principles, bidirectional steelman) packaged as a cross-agent SKILL.md for Codex, Claude Code, and Kimi Code.
- ai-avatar-skill — Purpose-driven AI avatars and profile pictures: Codex skill, ChatGPT prompts, examples, and crop previews.
- CHINESE-OCR
— Chinese scene-text detection and recognition; legacy project.
- Image2Katex — Image-to-LaTeX recognition for printed and handwritten formulas.
- DKT-TensorFlow — Deep knowledge tracing in TensorFlow.



