Zhining Zhang | 张芷宁

Hi! I am a first-year PhD student in EECS at UC Berkeley, affiliated with CHAI and BAIR. I am currently working with Professor Serina Chang.

I work on human-centered AI. Right now I am excited about human simulation and building AI that improves social welfare. If you'd like to chat about research or potential collaboration, please reach out!

I am deeply greatful to the amazing mentors, professors, and collaborators I've worked with, for their invaluable guidance and support along the way.


Experience
  • UC Berkeley
    UC Berkeley
    Ph.D. Student in Computer Science
    Sep. 2026 - Present
  • Peking University
    Peking University
    B.S. in Computer Science
    Sep. 2022 - Jun. 2026
  • Johns Hopkins University
    Johns Hopkins University
    Visiting Student
    June. 2024 - Sep. 2024
Honors & Awards
  • May Fourth Scholarship, Highest-level Scholarship @ Peking University
    2025
  • Merit Student, Peking University
    2025
  • Zhi-Class Scholarship
    2024
  • Peking University Scholarship
    2023
  • Award for Scientific Research, Peking University
    2023
Selected Publications (view all )
CocoaBench: Evaluating Unified Digital Agents in the Wild
CocoaBench: Evaluating Unified Digital Agents in the Wild

Shibo Hao*, Zhining Zhang*, Zhiqi Liang*, Tianyang Liu*, Yuheng Zha, Qiyue Gao, Jixuan Chen, Zilong Wang, Zhoujun Cheng, Haoxiang Zhang, Junli Wang, Hexi Jin, Boyuan Zheng, Kun Zhou, Yu Wang, Feng Yao, Licheng Liu, Yijiang Li, Zhifei Li, Zhengtao Han, Pracha Promthaw, Tommaso Cerruti, Xiaohan Fu, Ziqiao Ma, Jingbo Shang, Lianhui Qin, Julian McAuley, Eric P. Xing, Zhengzhong Liu, Rupesh Kumar Srivastava, Zhiting Hu (* equal contribution)

Conference on Language Modeling (COLM), 2026

CocoaBench evaluates unified digital agents on human-designed, long-horizon tasks that require flexible composition of vision, search, and coding, with instruction-only task specifications and automatic evaluation functions.

CocoaBench: Evaluating Unified Digital Agents in the Wild

Shibo Hao*, Zhining Zhang*, Zhiqi Liang*, Tianyang Liu*, Yuheng Zha, Qiyue Gao, Jixuan Chen, Zilong Wang, Zhoujun Cheng, Haoxiang Zhang, Junli Wang, Hexi Jin, Boyuan Zheng, Kun Zhou, Yu Wang, Feng Yao, Licheng Liu, Yijiang Li, Zhifei Li, Zhengtao Han, Pracha Promthaw, Tommaso Cerruti, Xiaohan Fu, Ziqiao Ma, Jingbo Shang, Lianhui Qin, Julian McAuley, Eric P. Xing, Zhengzhong Liu, Rupesh Kumar Srivastava, Zhiting Hu (* equal contribution)

Conference on Language Modeling (COLM), 2026

CocoaBench evaluates unified digital agents on human-designed, long-horizon tasks that require flexible composition of vision, search, and coding, with instruction-only task specifications and automatic evaluation functions.

Neural Synchrony Between Socially Interacting Language Models
Neural Synchrony Between Socially Interacting Language Models

Zhining Zhang, Wentao Zhu, Chi Han, Yizhou Wang, Heng Ji

International Conference on Learning Representations (ICLR), 2026

We introduce neural synchrony during social simulations as a proxy for analyzing the sociality of LLMs at the representational level, showing that it captures social engagement, temporal alignment, and correlations with social performance.

Neural Synchrony Between Socially Interacting Language Models

Zhining Zhang, Wentao Zhu, Chi Han, Yizhou Wang, Heng Ji

International Conference on Learning Representations (ICLR), 2026

We introduce neural synchrony during social simulations as a proxy for analyzing the sociality of LLMs at the representational level, showing that it captures social engagement, temporal alignment, and correlations with social performance.

AutoToM: Scaling Model-based Mental Inference via Automated Agent Modeling
AutoToM: Scaling Model-based Mental Inference via Automated Agent Modeling

Zhining Zhang*, Chuanyang Jin*, Mung Yao Jia*, Shunchi Zhang*, Tianmin Shu (* equal contribution)

Spotlight, Annual Conference on Neural Information Processing Systems (NeurIPS), 2025

We introduce AutoToM, an automated agent modeling method for scalable, robust, and interpretable mental inference. Leveraging an LLM as the backend, AutoToM combines the robustness of Bayesian models and the open-endedness of Language models, offering a scalable and interpretable approach to machine ToM.

AutoToM: Scaling Model-based Mental Inference via Automated Agent Modeling

Zhining Zhang*, Chuanyang Jin*, Mung Yao Jia*, Shunchi Zhang*, Tianmin Shu (* equal contribution)

Spotlight, Annual Conference on Neural Information Processing Systems (NeurIPS), 2025

We introduce AutoToM, an automated agent modeling method for scalable, robust, and interpretable mental inference. Leveraging an LLM as the backend, AutoToM combines the robustness of Bayesian models and the open-endedness of Language models, offering a scalable and interpretable approach to machine ToM.

Language models represent beliefs of self and others
Language models represent beliefs of self and others

Wentao Zhu, Zhining Zhang, Yizhou Wang

International Conference on Machine Learning (ICML) , 2024

We investigate belief representations in LMs: we discover that the belief status of characters in a story is linearly decodable from LM activations. We further propose a way to manipulate LMs through the activations to enhance their Theory of Mind performance.

Language models represent beliefs of self and others

Wentao Zhu, Zhining Zhang, Yizhou Wang

International Conference on Machine Learning (ICML) , 2024

We investigate belief representations in LMs: we discover that the belief status of characters in a story is linearly decodable from LM activations. We further propose a way to manipulate LMs through the activations to enhance their Theory of Mind performance.

All publications