|
Yiquan Wang (王一权)
AI for Science · 深度学习 · 计算生物学 · 生物信息学 · 蛋白质设计 · 分子动力学 · 三维基因组 · 虚拟细胞 · 表征学习 · 数学物理AI for Science · Deep Learning · Computational Biology · Bioinformatics · Protein Design · Molecular Dynamics · 3D Genome Folding · Virtual Cells · Representation Learning · Mathematical Physics
|
新疆大学 数学与应用数学(国家理科基地),乌鲁木齐
清华大学钱学森班 · 深圳零一学院 联合培养 · 零一学者
中国大学生自强之星(2025)
IEEE Biometrics Council · 中国化学会(CCS)学生会员
· ·
Xinjiang University — Mathematics & Applied Mathematics (National Base), Urumqi
Tsinghua TEEP · Shenzhen X-Institute — Joint Program · Zero One Scholar
China College Student Self-improvement Star (2025)
IEEE Biometrics Council · Chinese Chemical Society (CCS), Student Member
· ·
|
个人简介
王一权是新疆大学数学与应用数学专业(国家理科基地班)的本科生,同时也是清华大学钱学森班与深圳零一学院的联合培养零一学者。他的研究方向为科学智能(AI for Science),致力于解码生命系统的复杂性。
他的科研之路始于数学建模。本科早期,他在建模竞赛与科研项目中广泛探索,从图论到 AI 驱动的热浪灾害风险分析,论文列表里的跨领域条目正是那段旅程留下的脚印。数学建模训练了他把真实系统抽象为模型的直觉,也让他逐渐看清了自己真正的热情所在。蛋白质是研究生命系统的合适小模型,Anfinsen 法则指出氨基酸序列决定三维结构,问题定义清晰,序列、结构与功能注释数据完备,天然适合表征学习,加之个人兴趣,他将蛋白质确立为核心研究对象。如果说蛋白质是理解生命系统的小模型,基因组就是这套系统更完整的蓝图,博士生阶段他的研究将聚焦基因组,同时继续推进蛋白质组的工作。
在此基础上,他建立了两个独立而互补的研究方向:其一是数据驱动的表征学习与生物分子机制,利用深度学习解析复杂生物系统,研究涵盖蛋白质结构预测、蛋白质功能预测、蛋白质设计、分子动力学、多组学以及分子机制等多个生命科学问题;其二是基于数学物理的生物复杂性第一性原理,运用凝聚态物理、拓扑学、流形学习与微分几何推导生物系统的解析真理,独立于AI近似之外。他的核心愿景是将这两大方向融合,结合物理的归纳偏置与深度学习的表达能力,攻克一系列前沿科学难题,包括内在无序蛋白质(IDP)动态系综生成、液-液相分离(LLPS)与三维基因组折叠、靶向"不可成药"靶点的理性设计以及虚拟细胞的构建。通过这些突破,他最终旨在攻克复杂人类疾病,并走向对生命本质的根本理解。
研究成果发表于《Journal of Chemical Information and Modeling》《Physica D: Nonlinear Phenomena》《Information Sciences》等期刊,在IEEE BIBM、IEEE SMC等计算机主会上发表,并在ICLR、ICML、AAAI等会议的研讨会(workshop)中展示。欢迎就共同感兴趣的课题交流讨论。
教育经历
本科 · 数学与应用数学专业
数学与系统科学学院(数学物理研究所)
国家理科数学基地
课程与导师单位详情
主要课程:数学分析、高等代数、解析几何、偏微分方程、泛函分析等。
培养基地:国家理科基础学科研究和教学人才培养“数学”基地。参见数学基地简介
秦艳红:时任数学与系统科学学院/数学物理研究所,现物理科学与技术学院。
魏凯:新疆生物资源基因工程重点实验室·省部共建国家重点实验室培育基地,生命科学与技术学院。
联合培养
钱学森班
主要研究:ESRT · ORIC · SURF
参见钱班阶梯式培养方案 ↗
访问学生 · 袁文课题组
神经疾病研究所
研究方向:神经科学、生物化学、计算生物学、单细胞组学
访问学生 · 黄凯课题组
系统与物理生物学研究所
研究方向:计算生物物理,包括 IDP 动态系综生成、LLPS 与三维基因组折叠的解析,以及“不可成药”靶点的理性设计。
研究经历
代表性论文 (Featured Papers)
方向一:表征学习与生物分子机制 (Representation Learning & Biomolecular Mechanisms)
我目前的研究利用最前沿的深度学习技术来解析复杂的生物系统。我专注于通过引入新颖的归纳偏置和表征策略来突破传统生物信息学的局限。我的工作涵盖蛋白质结构预测、蛋白质设计和分子动力学,旨在为生命科学提供稳健、高通量的工具。
Wang, Y., Cai, M., et al. (2026). Sequence context decodes multi-state conformational heterogeneity from crystallographic B-factors. Science China Life Sciences, in press (earlier version: LMRL Workshop at ICLR 2026). 即使晶格堆积掩盖了部分信号,静态晶体结构中仍保留着蛋白质运动的线索。BFC 利用蛋白质语言模型提供的序列上下文,从 B 因子中恢复逐残基的构象异质性,与晶体学结构系综的异质性分布达到 0.83 的 Pearson 相关系数。分子对接、分子伴侣和抗体案例展示了校正信号在结构分析中的用途。
Wang, Y., Cai, M., et al. (2025). From Signal to Symphony: Exploring 2D Sequence Representations for Protein Function Prediction. Journal of Chemical Information and Modeling. [PDF] 氨基酸序列可以被读成乐谱:疏水性决定音高,分子量决定节奏。将序列转换为二维频谱图,为蛋白质功能预测提供了新的表征方式。基准评测与消融实验表明,主要预测信号来自一维到二维的结构转换,生物物理编码带来进一步增益;GFP 设计的计算概念验证则探索了这一表征的生成用途。
方向二:基于数学物理的生物复杂性第一性原理 (First-Principles of Biological Complexity)
虽然深度学习提供了强大的近似能力,但我的终极目标是找到生物系统的解析基本真理。我正在从纯数据驱动方法转向物理启发理论,旨在利用凝聚态物理、拓扑学、流形学习和微分几何来描述蛋白质的能量景观和生命的热力学本质。
方向三:跨学科探索与协同效应 (Interdisciplinary Explorations & Collaborative Synergies)
我的研究之旅源于广泛的好奇心。我与语言学、计算机视觉、密码学和运筹学等领域的专家合作。这些多元化的经历赋予了我独特的工具箱——使我能够跨领域迁移方法论,并从正交视角解决问题。
Wang, Y., Zai, J., Liu, Z., Chen, J., & Sabir, E.* (2026). Hamiltonian Connectivity in k-Ary n-Cubes under a Region-Based Fault Model. Information Sciences. 成片故障对网络可靠性的影响与孤立故障不同。区域故障模型将损伤表示为具有空间约束的连通子图,并给出奇数 k ≥ 3、n ≥ 2 时 k 元 n 立方体保持哈密尔顿连通的充分条件。构造性的自适应算法支撑了证明,实验则指出故障簇之间的空间分离是网络韧性的关键因素。
Wang, Y.*, et al. (2026). AI Driven Discovery of Bio Ecological Mediation in Cascading Heatwave Risks. Physics and Chemistry of the Earth. 热浪风险跨越多个学科传播,相关证据却散落在数千篇研究中。HeDA 将 8,365 篇文献组织为含 34,933 个实体、42,890 条有向关系的知识图谱。留出推理测试显示了图谱增强的收益,重建的风险拓扑则突出农业与人体健康的结构中介作用,连接热应激与更广泛的社会经济损失。
查看全部论文列表 | Google Scholar主页
完整科研项目、研学、实习、竞赛与审稿经历,见 经历页(Experience)。
微信公众号
|
biomath
聚焦 AI for Science 与计算生物学前沿:从蛋白质设计、分子与宏观演化,到多尺度理化模拟的交叉科学探索。
🧬 大模型与蛋白质设计
🌿 分子演化与宏观进化生物学
🧪 计算化学与多组学探索
⚛️ 物理生物学与前沿交叉科学
欢迎扫码关注,交流探讨!
|
简历下载.
Personal Profile
Yiquan Wang is an undergraduate student in Mathematics and Applied Mathematics at the National Base for Research and Teaching Talents at Xinjiang University, and a joint-training scholar in the Tsien Excellence in Engineering Program at Tsinghua University & Shenzhen X-Institute. His research focuses on AI for Science, driven by an enduring pursuit to decode the complexity of living systems.
His research journey began with mathematical modeling. In his early undergraduate years, he explored widely through modeling competitions and research training programs, working on graph theory and AI-driven heatwave risk analysis, directions that left their footprints across the interdisciplinary entries in his publication list. Mathematical modeling trained him to abstract real-world systems into models, and gradually revealed where his true passion lay. Proteins are a suitable small model for studying living systems. Anfinsen's dogma states that the amino acid sequence determines the three-dimensional structure, which makes the problem well-posed, and sequence, structure, and functional annotation data are abundantly available, a natural fit for representation learning; together with his personal interest, this led him to settle on proteins as his core research object. If proteins are the small model for understanding living systems, the genome is the blueprint of the same system in full. In his PhD stage, his research will focus on genomics while continuing his work on proteomics.
He has since established two distinct yet complementary research directions. The first centers on data-driven Representation Learning and Biomolecular Mechanisms, applying deep learning to a broad range of life science problems spanning protein structure prediction, protein function prediction, protein design, molecular dynamics, multi-omics, and molecular mechanisms. The second pursues First-Principles of Biological Complexity Based on Mathematical Physics, deriving analytical ground truths of biological systems through condensed matter physics, topology, manifold learning, and differential geometry, independent of AI approximations. His overarching goal is to synergize these two directions, combining the inductive biases of physics with the expressive power of deep learning to tackle a constellation of formidable open problems, including Intrinsically Disordered Protein (IDP) dynamic ensemble generation, Liquid-Liquid Phase Separation (LLPS) and 3D genome folding, rational design for "undruggable" targets, and the construction of Virtual Cells. Through these advances, he ultimately aims to confront complex human diseases and move closer to a fundamental understanding of life itself.
His work has been published in journals including the Journal of Chemical Information and Modeling, Physica D: Nonlinear Phenomena, and Information Sciences, presented at main conferences including IEEE BIBM and IEEE SMC, and presented at workshops of ICLR, ICML, and AAAI. Discussions and collaborations on topics of shared interest are always welcome.
Education
Undergraduate · Mathematics and Applied Mathematics
College of Mathematics and System Sciences (Institute of Mathematics and Physics)
National Mathematics Base
Coursework & affiliation details
Main Courses: Mathematical Analysis, Advanced Algebra, Analytical Geometry, Partial Differential Equations, Functional Analysis, etc.
Program: National Base for Research and Teaching Talents in Basic Sciences "Mathematics". Mathematics Base Introduction
Yan-Hong Qin: Then at the College of Mathematics and System Sciences / Institute of Mathematics and Physics; now at the School of Physical Science and Technology.
Kai Wei: Xinjiang Key Laboratory of Biological Resources and Genetic Engineering, State Key Laboratory Incubation Base co-built by the Province and Ministry, College of Life Science and Technology.
Visiting Student · Wen Yuan Research Group
Institute of Neurological and Psychiatric Disorders
Research Interests: Neuroscience, biochemistry, computational biology, single-cell omics
Visiting Student · Kai Huang Research Group
Institute of Systems and Physical Biology
Research Interests: Computational biophysics, including IDP dynamic ensemble generation; decoding LLPS & 3D genome folding; rational design for "undruggable" targets.
Research Experience
Featured Papers
Direction 1: Representation Learning & Biomolecular Mechanisms
My current research leverages state-of-the-art Deep Learning techniques to decipher complex biological systems. I focus on overcoming the limitations of traditional bioinformatics by introducing novel inductive biases and representation strategies. My work spans protein structure prediction, protein design, and molecular dynamics, aiming to provide robust, high-throughput tools for the life sciences.
Wang, Y., Cai, M., et al. (2026). Sequence context decodes multi-state conformational heterogeneity from crystallographic B-factors. Science China Life Sciences, in press (earlier version: LMRL Workshop at ICLR 2026). Static crystal structures retain clues to protein motion even when lattice packing obscures them. BFC uses protein-language-model sequence context to recover residue-wise conformational heterogeneity from B-factors, matching crystallographic ensemble profiles with a Pearson correlation of 0.83. Docking, chaperone, and antibody case studies illustrate how the corrected signal can support structural analysis.
Wang, Y., Cai, M., et al. (2025). From Signal to Symphony: Exploring 2D Sequence Representations for Protein Function Prediction. Journal of Chemical Information and Modeling. [PDF] Amino acid sequences can be read as musical scores: hydrophobicity sets pitch and molecular weight sets rhythm. Converting these sequences into two-dimensional spectrograms provides useful representations for protein function prediction. Benchmarking and ablations show that much of the predictive signal comes from the 1D-to-2D transformation itself, with biophysical encoding providing a further gain; a computational proof-of-concept extends the representation to GFP design.
Direction 2: First-Principles of Biological Complexity Based on Mathematical Physics
While Deep Learning provides powerful approximations, my ultimate goal is to find the analytical ground truths of biological systems. I am transitioning from purely data-driven approaches to physics-informed theories. I aim to use Condensed Matter Physics, Topology, Manifold Learning, and Differential Geometry to describe the energy landscapes of proteins and the thermodynamic essence of life.
Wang, Y. (2026). Descriptive power and predictive limits of a discrete Hasimoto–DNLS model of protein backbone structure. Physica D: Nonlinear Phenomena. [arXiv] An exact decomposition of the Hasimoto–DNLS effective potential separates chirality from the contributions of local backbone geometry. Analysis across 856 proteins shows why this compact representation describes folded structures yet fails to predict native folds through a local real-potential reduction. The map is a kinematic identity, and its dispersion residual provides a geometric marker of near-integrable α-helices.
Wang, Y. (2026). Spectral analysis of protein backbone geometry reveals abrupt helix–coil boundaries. The European Physical Journal E. [arXiv] Helices and coils leave distinct spectral signatures in backbone geometry. Mapping 1,986 protein structures through the discrete Hasimoto transform reveals abrupt, directionally asymmetric boundaries between low-entropy helices and broadband coils. Pointwise integrability and windowed spectral probes capture complementary features; the analysis also explains boundary assignment ambiguity and the spatial–spectral trade-off that limits windowed measurements.
Direction 3: Interdisciplinary Explorations & Collaborative Synergies
My research journey is fueled by broad curiosity. I have collaborated with experts in linguistics, computer vision, cryptography, and operations research. These diverse experiences have equipped me with a unique toolkit—allowing me to transfer methodologies and tackle problems from orthogonal perspectives.
Wang, Y., Zai, J., Liu, Z., Chen, J., & Sabir, E.* (2026). Hamiltonian Connectivity in k-Ary n-Cubes under a Region-Based Fault Model. Information Sciences. Clustered failures challenge network reliability differently from isolated faults. The Region-Based Fault model represents damage as spatially constrained connected subgraphs and establishes conditions for Hamiltonian connectivity in k-ary n-cubes with odd k ≥ 3 and n ≥ 2. A constructive adaptive algorithm supports the proof, while experiments identify separation between fault clusters as a key factor in resilience.
Wang, Y.*, et al. (2026). AI Driven Discovery of Bio Ecological Mediation in Cascading Heatwave Risks. Physics and Chemistry of the Earth. Heatwave risks propagate across disciplinary boundaries, but the evidence is scattered across thousands of studies. HeDA organizes 8,365 publications into a knowledge graph with 34,933 entities and 42,890 directed relationships. Held-out reasoning tests show benefits from graph augmentation, while the reconstructed topology highlights agriculture and human health as structural mediators linking thermal stress to wider socioeconomic losses.
View full publication list | Google Scholar Profile
The full record of research projects, study programs, internships, competitions, and reviewing lives on the Experience page.
WeChat Official Account
|
biomath
Exploring the frontiers of AI for Science and Computational Biology: from protein design, molecular and macroevolution, to multi-scale physicochemical simulations.
🧬 Large Models & Protein Design
🌿 Molecular Evolution & Macroevolutionary Biology
🧪 Computational Chemistry & Multi-omics
⚛️ Physical Biology & Frontier Interdisciplinary Science
Scan the QR code to follow!
|
Curriculum Vitae Download.
|