About Me

😊 Hi, I am Ziyi Yang, a Ph.D. candidate in Organic Chemistry at Tsinghua University. My work lies at the intersection of computational protein design and chemical protein synthesis, with the goal of designing and synthesizing functional proteins, including therapeutics, catalysts, materials, and beyond, that go beyond what evolution on Earth has explored.

Download CV GitHub

Research

I have a background in chemical biology and chemical protein synthesis. Since 2023, I have worked in the Lei Liu group at the Department of Chemistry, Tsinghua University, where I have focused on developing new chemical reactions and strategies to enable the total synthesis of larger and more diverse proteins.

As our synthetic capabilities continue to grow rapidly, one fundamental question remains: Do we really need chemical protein synthesis? Given that biological technologies such as recombinant expression can already produce most existing proteins, chemical synthesis is most compelling when it enables access to noncanonical proteins that cannot be produced by biological systems. These include proteins with precise post-translational modifications, artificial functional groups, or noncanonical amino acids, namely, proteins that go beyond the 20 amino acids selected by evolution.

For natural proteins, evolution has already provided clues about which sequences are foldable and functional, making the synthetic targets relatively clear. For noncanonical proteins, however, we do not yet have such guidance. We may say that chemical synthesis gives us the possibility to “create a new nature beside the old” — to borrow R. B. Woodward’s phrase — but we can only do so if we have a blueprint that tells us what to synthesize.

The emergence of AI offers an opportunity to transfer knowledge from natural proteins and small molecules to the prediction and design of noncanonical proteins. Therefore, since April 2025, I have started exploring computational methods for noncanonical amino acid protein design.

My previous research has focused on synthetic data and representation learning, as summarized below. I am currently interested in reinforcement learning methods that incorporate feedback from wet-lab experiments to improve generative models. I am also exploring long-term scientific questions, such as: (1) Is homochirality an evolutionary accident, or is it dictated by a fundamental principle, namely that stable folding of heterochiral proteins is extremely rare? (2) Are there more folding motifs in the universe if we introduce more noncanonical amino acids to increase structural complexity? (3) Do valuable functional molecules exist beyond the chemical space adopted by life on Earth?

With the power of AI and chemical protein synthesis, I believe we can begin to answer these questions.

Here is a summary of my previous research.

Computational protein design

  • Cyclic peptide binders. I build large-scale cyclic-peptide–receptor structure datasets by mimicking inter-chain interactions with intra-protein interfaces. This line of work supports model scaling for cyclic peptide binder design, and our synthetic data have been used by several teams.
  • Mirror-image peptide binders. I develop chirality-aware generative modeling methods for heterochiral design, including axial-vector and pseudoscalar features that help models learn and control molecular handedness.

Chemical protein synthesis

  • Glycosylation-assisted protein folding. I study how glycosylation can suppress aggregation during folding and enable the synthesis of proteins that are otherwise difficult to fold.
  • Triple-split inteins. I work on split intein systems for the total synthesis of ultra-large proteins (>100 kDa).

Selected Publications

  • Cyclic peptide binder design. CPSea1 was presented at NeurIPS 2025, and CPSea2 is currently under review. Dataset: CPSea repo.
  • Chirality-aware generative modeling. Mirror-image peptide binder design with explicit chirality features was accepted to ICML 2026. Preprint: arXiv:2602.20176.
  • Glycosylation-assisted folding. Published in Angewandte Chemie in 2024 with Tian Wang, Wenjun Shi, and collaborators.
  • Ultra-large protein synthesis. Ongoing manuscript on triple-split inteins for the total synthesis of proteins larger than 100 kDa. I mainly focus on this project in the first two years of my PhD.

Experience

  • ByteDance AI Drug Discovery, Intern, peptide drug design, 2026.04–2026.06.
  • Yanyan Lan Lab, Institute for AI Industry Research (AIR), Tsinghua University, Intern, protein design with synthetic data and generative models, 2025.03–present.
  • Jishen Zheng Lab, University of Science and Technology of China, Summer visitor, directed evolution of split inteins via phage display, 2023.06–2023.07.
  • Peilong Lu Lab, Westlake University, Summer visitor, mirror-image peptide binder design via the RIF pipeline, 2022.07–2023.08.

Education

  • Ph.D. in Organic Chemistry, Tsinghua University, 2023–present. Expected graduation: June 2028.
  • B.S. in Chemical Biology, Tsinghua University, 2019–2023.

Contact

I am happy to chat about AI, science, and visions for the future.

Email: yang-zy19@outlook.com yangzy23@mails.tsinghua.edu.cn