Shengzhuang Chen

sheng<dot>chen17[at]imperial<dot>ac<dot>uk
Shengzhuang Chen

I am a final-year PhD student at the Imperial Frontier AI Research Lab, supervised by Prof. Jonathan Richard Schwarz and Prof. Alessandra Russo. I am also a Lead Research Scientist at Thomson Reuters Foundational Research, where I work on mid- and post-training of LLMs at scale for enterprise applications, including agentic and tool-use systems.


My research focuses on the data efficiency and generalization of LLM post-training, using meta-learning, sparsity, and data-centric methods. Previously, I was a visiting student at Harvard University; a PhD student at City University of Hong Kong with Prof. Ying Wei, supported by the Hong Kong PhD Fellowship (HKPFS); and an MEng graduate of Imperial College London in Electrical and Electronic Engineering with first-class honours (top 10% of the cohort). I have published multiple first-author papers at ICML, NeurIPS, ICLR, ACL (including an Oral), and TMLR. If you are interested in my work or would like to collaborate, feel free to reach out! 😊

News

Aug 2026 We released Thomson, a family of frontier models built by continual learning on open-weight checkpoints for high-stakes professional work — I am honoured to be a primary author of the technical report. We also open-sourced the 35B Thomson-1.0-Small on Hugging Face.
Dec 2025 ADMIRE-BayesOpt accepted to TMLR.
Oct 2025 Invited talk on Data-Centric ML for LLMs at the Allen Institute for AI (AI2).
Jul 2025 Presented SIMoE as an oral presentation at ACL 2025 in Vienna.
Jun 2025 Invited talk on Sparse Interpolated Mixture-of-Experts for LLM Upcycling at Qingke AI.
Jan 2025 Two papers accepted — CLDyB at ICLR 2025 and SIMoE at ACL 2025.
Jan 2025 Transferred PhD to Imperial College London.
Jan 2025 Started research scientist internship at Thomson Reuters Foundational Research.
Sep 2024 Paper on Learning Where to Edit Vision Transformers accepted to NeurIPS 2024.
May 2024 Paper on Sparse Interpolated Experts for Few-Shot Generalization accepted to ICML 2024.
May 2024 Started visiting research at Harvard Medical School, hosted by Prof. Marinka Zitnik.
Sep 2023 Paper on Secure OOD Task Generalization with EBMs accepted to NeurIPS 2023.
Sep 2022 Awarded the Hong Kong PhD Fellowship (HKPFS).

Selected Publications

Technical Reports

Conference & Journal Publications

SMAT overview
Unleashing the Power of Meta-tuning for Few-shot Generalization Through Sparse Interpolated Experts
S. Chen, J. Tack, Y. Yang, Y. W. Teh, J. R. Schwarz, Y. Wei
ICML 2024
Abstract
We propose SMAT, a meta-tuning approach that constructs sparse interpolated experts by blending pre-trained and fine-tuned model parameters. A learned sparse routing mechanism selects and combines these experts at test time, yielding strong few-shot generalization across diverse tasks with improved parameter efficiency.
meta-learning few-shot learning mixture-of-experts
EBML overview
Secure Out-of-Distribution Task Generalization with Energy-Based Models
S. Chen, L.-K. Huang, J. R. Schwarz, Y. Du, Y. Wei
NeurIPS 2023
Abstract
We present a unified framework that leverages energy-based models to jointly detect and adapt to out-of-distribution tasks in few-shot settings. The energy function provides a principled score for OOD detection while also guiding task-conditional adaptation, enabling reliable generalization across diverse domains.
meta-learning few-shot learning

Research Experience

Lead Research Scientist, Foundational Research Jan 2025 – Present
Thomson Reuters Ltd
Continual pre-training and post-training of hundred-billion-parameter foundation models.
Visiting Student May – Jul 2024
Harvard Medical School, Harvard University
Instruction-tuning for LLMs, with an emphasis on improving downstream cross-task generalization performance. Hosted by Prof. Marinka Zitnik.

Invited Talks

Data-Centric ML for LLMs Oct 2025
Allen Institute for AI (AI2)
Sparse Interpolated Mixture-of-Experts for LLM Upcycling Jun 2025
Qingke AI

Education

Ph.D. Candidate in Computer Science 2025 – Exp. EOY 2026
Imperial College London
Supervisors: Prof. Jonathan Richard Schwarz and Prof. Alessandra Russo.
Ph.D. Candidate in Computer Science 2022 – 2025
City University of Hong Kong
Supervisor: Prof. Ying Wei. Transferred to Imperial College London in 2025.
M.Eng. in Electrical and Electronic Engineering 2017 – 2021
Imperial College London
First class honours; top 10% of the cohort.

Awards & Honours

PhD Scholarship, TR–Imperial Frontier AI Lab 2025
Thomson Reuters & Imperial College London — Fully funded PhD studentship.
Research Tuition Scholarship 2024
City University of Hong Kong — Awarded for outstanding research contributions and academic performance.
HKPFS Academic Excellence Award 2024
City University of Hong Kong — Recognizes exceptional academic achievements among HKPFS recipients.
Hong Kong PhD Fellowship (HKPFS) 2022
Research Grants Council of Hong Kong — Highly competitive fellowship supporting outstanding doctoral students (<5% acceptance rate).
Nicholas Battersby Prize 2021
Imperial College London — Best Master Thesis in Analogue Electronics.
Dean's List for Academic Excellence 2017, 2018, 2021
Imperial College London — Top 10% of the cohort for exceptional academic performance.

Academic Service

Conference Reviewer
ICML, ICLR, NeurIPS, ACL, TMLR
Teaching Assistant 2022 – 2025
CS2402 Computational Probability Modeling (2024–2025) · CS5491 Artificial Intelligence (2022–2023), City University of Hong Kong