|
Experience
科研项目
蛋白质声化将氨基酸序列转换为二维频谱图,用于功能预测。包含 12 类、18,000 条序列的基准支撑了表征研究,以及 GFP 设计的计算概念验证。
摘要从蛋白质一级序列预测其功能是计算生物学中的一个基本挑战。虽然深度学习已经取得了显著成就,但序列数据的最优表示方式仍然是一个开放问题。本研究探索了蛋白质声化——将氨基酸序列转换为二维频谱图——作为该任务的表示方法。为了促进这一研究,我们开发了一个包含18,000条序列的基准数据集,涵盖12个功能多样的蛋白质类别。我们的系统评估表明,从一维序列到二维频谱图的结构转换可能是模型预测性能的关键因素。这一观察得到了消融研究的支持,其中仅使用频谱图中的视觉或声学特征的模型都表现出了有效的独立性能,表明该表示本身是这种能力的关键来源。例如,使用没有明确生物物理意义的声化映射的模型达到了81.08%的准确率,而我们的生物物理信息模型达到了84.00%,表明这样的领域知识可能提供了适度的性能提升。当在我们的数据集上从头开始训练时,我们的融合模型的性能与标准Transformer架构(如ESM-2和ProtBERT)相当或略优,表明其在这一特定背景下的数据效率潜力。该模型的泛化潜力进一步得到了其在外部CARE酶分类基准上的性能支持,其中它达到了90.44%的准确率。最后,作为概念验证,我们探索了我们的编码在指导扩散模型生成新颖GFP变体中的效用,这些变体使用计算方法进行了结构可行性评估。我们的工作提供了证据,表明声化在这一背景下的效用可能主要源于其表示结构,为生物序列的特征工程提供了一个视角。
关键词:蛋白质功能预测、声化、生物信号处理、序列表示、生成蛋白质设计
项目地址: GitHub/Symphony_of_Fate · 指导老师:魏凯教授
融合全基因组测序中的拷贝数变异与临床特征,探索用于阿尔茨海默病风险评估的条件扩散框架。
摘要本项目旨在将全基因组测序数据中的基因拷贝数变异(Copy Number Variation, CNV)特征与代谢指标等多维临床数据相融合。利用扩散模型在处理高维模拟数据和蛋白质表型预测中的成功经验,我们将构建一个由CNV特征编码、基因组区域注意力和条件U-Net扩散模块组成的综合框架。这将模拟CNV在基因组中的分布变化和进化过程,分析CNV在调控阿尔茨海默病通路中的具体作用机制,最终提高疾病风险评估和早期干预的准确性。
关键词:Diffusion; Copy Number Variation; WGS; Alzheimer's disease
指导老师:魏凯教授
热浪风险跨越多个学科传播,相关证据却散落在数千篇研究中。HeDA 将 8,365 篇文献组织为含 34,933 个实体、42,890 条有向关系的知识图谱。留出推理测试显示了图谱增强的收益,重建的风险拓扑则突出农业与人体健康的结构中介作用,连接热应激与更广泛的社会经济损失。
摘要复合热浪会在相互关联的物理与人类系统中触发级联故障,而分散的学科知识使系统性风险拓扑难以被完整描绘。HeDA 从 8,365 篇学术文献构建知识图谱,利用按时间划分的建图语料(截至 2024 年,5,739 篇)组织 34,933 个实体与 42,890 条有向关系。在按时间留出的多跳推理基准(2025—2026 年,3,892 个问题)上,图谱增强相对四个独立基础模型带来 8.1%—9.8% 的准确率提升。拓扑分析识别出连接热应激与系统性社会经济损失的生物系统结构中介作用,并发现部门之间的潜在功能耦合,例如热应激下电网故障与急诊医疗容量饱和的同步化。这些结果为面向相互关联系统的韧性适应策略提供了结构依据。
关键词:知识图谱、智能代理、热浪风险分析、科学发现、气候适应、多层风险传播
项目地址: GitHub/heatwave · 指导老师:葛咏研究员
成片故障对网络可靠性的影响与孤立故障不同。区域故障模型将损伤表示为具有空间约束的连通子图,并给出奇数 k ≥ 3、n ≥ 2 时 k 元 n 立方体保持哈密尔顿连通的充分条件。构造性的自适应算法支撑了证明,实验则指出故障簇之间的空间分离是网络韧性的关键因素。
摘要$k$元$n$立方体($Q_n^k$)是支持现代AI和HPC工作负载的大规模计算系统的关键拓扑,其中容错能力至关重要。传统的故障模型假设故障是独立和随机的,会产生不切实际的弹性估计,因为现实世界的故障在空间上是相关的,表现为拓扑集群。本文引入了基于区域的故障(RBF)模型,这是一个新的范式,通过直接建模这种空间相关性来解决这一差距。我们的主要贡献是证明了对于奇数$k \geq 3$和$n \geq 2$,$Q_n^k$在充分的RBF条件下保持哈密尔顿连通性——这是无死锁路由和高效任务调度的关键特性。我们提出了一个构造性算法,通过利用自适应分解策略来找到哈密尔顿路径。实验分析表明,我们的方法显著增强了容错能力,并且在远超其保守理论保证的范围内保持稳健。这项工作为在故障聚类普遍存在的系统中维持连通性提供了一个实用的、高性能的解决方案。
关键词:$k$元$n$立方体、容错嵌入、哈密尔顿连通、聚类故障、互连网络
项目地址: GitHub/Hamiltonian_Path · 指导老师:依明江·沙比尔副教授
探索以个性化 AI 数字分身支持老年人的亲情陪伴与代际交流,同时结合中小学生的人工智能实践教育。
摘要本项目以人工智能数字分身技术为核心,旨在缓解我国老龄化背景下老年人亲情陪伴缺失的问题,同时弥合老年人数字鸿沟,提升中小学生的人工智能实践能力。项目由深圳零一学院联合字节跳动等机构实施,分为准备、执行和推广三个阶段,计划通过学生为老年人定制个性化数字分身,增进代际交流,探索AI大模型在社会公益领域的应用场景。项目不仅在精神层面为老年人提供陪伴与心理疏导,还开辟了中国特色养老新路径,同时推动人工智能教育与社会创新发展。通过多方协作、风险管理和媒体宣传,项目将实现从试点到全国推广的目标,具有重要的社会、教育和技术意义。
关键词:人工智能; 社会创新
指导老师:汤敏(国务院原参事)
Application of Artificial Intelligence in Whole Genome Selection Breeding.
Urban Big Data, Multimodal Fusion and Spatial Intelligence Analysis (第二作者)。
Forging a Strong Sense of Community for the Chinese Nation Exhibition Project (参与者)。
学术参与
审稿经历
会议研讨会
- ICLR · LMRL2026
- ICLR · FM4Science2026
- ICML · AI4Science2026
- NeurIPS · MATH-AI2026 · 2025
- NeurIPS · AI for Science2026 · 2025
- ICML · AI for Math2025
- ICLR · AI for Nucleic Acids2025
- ICLR · GenAI Watermarking2025
研学经历- 2026.04
- 2025.07
- 2025.07–2025.09
- 2025.01
- 2024.09–2024.12
- 2024.08
- 2024.08
- 2024.06–2027.06
清华大学钱学森力学班 · 深圳零一学院(Shenzhen X-Institute)
- 2024.03–2024.06
实习经历
清华大学北京生物结构前沿研究中心 2025.06-present:我开发了多肽结构数据库,从事蛋白质、多肽设计等工作。指导老师:Prof. Yafei Yuan、Prof. Tong Wang。
华为昇思Mindspore社区联合中科院软件研究所开源实习 2024.09-2025.3: 我利用机器学习、人工智能等技术,实现了基于VGG19的波洛克风格迁移画分形和湍流特征提取及NFT标签生成。目前已被华为公众号报道
玻色量子"星火人"社区实习生 2024.09-present
竞赛荣誉
全国大学生生命科学竞赛 国家二等奖, 2026
“挑战杯”大学生课外学术科技作品竞赛 自治区二等奖, 2026
全国大学生电子商务“创新、创意及创业”挑战赛 自治区二等奖, 2026
联合国全球领导力与ESG发展中心 × 深圳零一学院 "AI助力一老一小"创新挑战赛 最佳创意奖, 2026
全国大学生数学建模竞赛(CUMCM) 国家级二等奖, 2025.11
智慧化学城市挑战赛(新疆站) 二等奖, 2025
合成生物学创新赛 银奖, 2025 参见链接
The Mathematical Contest in Modeling (MCM, 美国大学生数学建模竞赛) Honorable Mention, 2025.5
2024阿里云天池大学生竞赛全国总决赛第17名, 2024.10参见链接
2024年第十四届APMCM亚太地区大学生数学建模竞赛国家级三等奖, 2024.08
"阿尔法蛋杯"2024年全国业余围棋棋王争霸赛暨"商旅运河杯"城市围棋赛竞赛第15名, 2024.07
新疆青少年业余围棋段位赛第53名, 2024.05
湖南省迎春杯围棋赛第七名, 2024.02
全国青少年智力运动大会围棋赛项第九名, 2024.02
新疆大学漏洞报送荣誉, 2023.10
2023年新疆"天山固网杯"网络安全技能竞赛第七名, 2023.10
深圳零一学院简介
深圳零一学院缘起于清华大学"学堂计划"钱学森力学班(简称"清华钱班")。清华钱班创办于2009 年,是"清华学堂人才培养计划"暨国家"基础学科拔尖学生培养试验计划"66个试点项目中,唯一不是定位于单一学科,而是工科基础(或力学与工程技术所有学科交叉创新)的试验班。其使命是:发掘和培养有志于通过科技改变世界、造福人类的创新型人才,探索未来创新人才的培养模式,回答"钱学森之问"。
Research Projects
Protein sonification turns amino acid sequences into two-dimensional spectrograms for function prediction. A benchmark of 18,000 sequences across 12 classes supports the representation study and a computational GFP design proof-of-concept.
AbstractPredicting protein function from its primary sequence is a fundamental challenge in computational biology. While deep learning has excelled, the optimal representation of sequence data remains an open question. This study explores protein sonification---the conversion of amino acid sequences into 2D spectrograms---as a representation for this task. To facilitate this investigation, we developed a benchmark dataset of 18,000 sequences spanning 12 functionally diverse protein classes. Our systematic evaluation suggests that the structural transformation from a 1D sequence to a 2D spectrogram may be a key contributor to the model's predictive performance. This observation is supported by ablation studies where models using either purely visual or acoustic features from the spectrogram demonstrated effective standalone performance, suggesting that the representation itself is a key source of this capability. For instance, a model using a sonification map without explicit biophysical meaning achieved 81.08% accuracy, while our biophysically-informed model reached 84.00%, indicating that such domain knowledge may offer a modest performance benefit. When trained from scratch on our dataset, our fusion model achieved performance comparable to or slightly exceeding that of standard transformer architectures like ESM-2 and ProtBERT, suggesting its potential for data efficiency in this specific context. The model's potential for generalizability was further supported by its performance on the external CARE enzyme classification benchmark, where it achieved 90.44% accuracy. Finally, as a proof-of-concept, we explore the utility of our encoding to guide a diffusion model in generating novel GFP variants, which were assessed for structural viability using computational methods. Our work provides evidence suggesting that the utility of sonification in this context may stem largely from its representational structure, offering a perspective on feature engineering for biological sequences.
Keywords: Protein Function Prediction, Sonification, Biological Signal Processing, Sequence Representation, Generative Protein Design
Project Address: GitHub/Symphony_of_Fate · Supervised by Prof. Kai Wei.
Copy number variation from whole-genome sequencing is combined with clinical features to develop a conditional diffusion framework for Alzheimer’s disease risk assessment.
AbstractThis project aims to integrate copy number variation (CNV) features from whole-genome sequencing data with multi-dimensional clinical data such as metabolic indicators. Leveraging the success of diffusion models in processing high-dimensional simulation data and protein phenotype prediction, we will construct a comprehensive framework consisting of CNV feature encoding, genomic region attention, and conditional U-Net diffusion modules. This will simulate CNV distribution changes and evolutionary processes in the genome, analyze the specific role of CNV in regulating Alzheimer's disease pathways, and ultimately improve disease risk assessment and early intervention accuracy.
Keywords: Diffusion; Copy Number Variation; WGS; Alzheimer's disease
Supervised by Prof. Kai Wei.
Heatwave risks propagate across disciplinary boundaries, but the evidence is scattered across thousands of studies. HeDA organizes 8,365 publications into a knowledge graph with 34,933 entities and 42,890 directed relationships. Held-out reasoning tests show benefits from graph augmentation, while the reconstructed topology highlights agriculture and human health as structural mediators linking thermal stress to wider socioeconomic losses.
AbstractCompound heatwaves increasingly trigger cascading failures that propagate through interconnected physical and human systems, yet fragmented disciplinary knowledge makes these systemic risk topologies difficult to map. HeDA constructs a knowledge graph from 8,365 academic publications, structuring 34,933 entities and 42,890 directed relationships from a temporally partitioned corpus (through 2024, n = 5,739). On a temporally held-out multi-hop reasoning benchmark (2025–2026, n = 3,892), graph augmentation improves accuracy by 8.1%–9.8% over four standalone foundation-model baselines. Topological analysis identifies biological systems as structural mediators linking thermal stress to systemic socioeconomic losses. The framework also identifies latent functional couplings between sectors, including the synchronization of power-grid failures and emergency medical capacity saturation under heat stress. These findings provide a structural basis for adaptation strategies that address resilience across interconnected systems.
Keywords: Knowledge Graph, Intelligent Agents, Heatwave Risk Analysis, Scientific Discovery, Climate Adaptation, Multi-layer Risk Propagation
Project Address: GitHub/heatwave · Supervised by Researcher Yong Ge.
Clustered failures challenge network reliability differently from isolated faults. The Region-Based Fault model represents damage as spatially constrained connected subgraphs and establishes conditions for Hamiltonian connectivity in k-ary n-cubes with odd k ≥ 3 and n ≥ 2. A constructive adaptive algorithm supports the proof, while experiments identify separation between fault clusters as a key factor in resilience.
AbstractThe $k$-ary $n$-cube ($Q_n^k$) is a critical topology for large-scale computing systems powering modern AI and HPC workloads, where fault tolerance is paramount. Traditional fault models, which assume faults are independent and random, yield unrealistic resilience estimates because real-world failures are spatially correlated, manifesting as topological clusters. This paper introduces the Region-Based Fault (RBF) model, a new paradigm that addresses this gap by directly modeling this spatial correlation. Our primary contribution is a proof that for odd $k \geq 3$ and $n \geq 2$, the $Q_n^k$ remains Hamiltonian-connected—a property vital for deadlock-free routing and efficient task scheduling—under a sufficient set of RBF conditions. We present a constructive algorithm that finds a Hamiltonian path by leveraging an adaptive decomposition strategy. Experimental analysis demonstrates that our approach significantly enhances fault tolerance and remains robust far beyond its conservative theoretical guarantees. The resulting framework provides a practical, high-performance solution for maintaining connectivity in systems where fault clustering is prevalent.
Keywords: $k$-ary $n$-cubes, fault-tolerant embedding, Hamiltonian-connected, clustered faults, interconnection networks.
Project Address: GitHub/Hamiltonian_Path · Supervised by Prof. Eminjan Sabir.
Personalized AI digital twins are explored as a way to support older adults’ family companionship and intergenerational communication, alongside practical AI education for school students.
AbstractThis project, centered on artificial intelligence digital twin technology, aims to alleviate the lack of family companionship for the elderly in China's aging society, bridge the digital divide for older adults, and enhance primary and secondary school students' artificial intelligence practical capabilities. Implemented by Shenzhen X-Institute in collaboration with ByteDance and other organizations, the project is divided into preparation, implementation, and promotion phases. It plans to facilitate students in customizing personalized digital twins for the elderly, enhancing intergenerational communication, and exploring application scenarios of AI large models in social welfare. The project not only provides companionship and psychological support for the elderly at a spiritual level but also opens up a new path for Chinese-style elderly care while promoting artificial intelligence education and social innovation development. Through multi-party collaboration, risk management, and media promotion, the project will achieve the goal of progressing from pilot programs to nationwide promotion, with significant social, educational, and technological implications.
Keywords: Artificial Intelligence; Social Innovation
Supervised by Min Tang (Former Counselor of the State Council).
Application of Artificial Intelligence in Whole Genome Selection Breeding.
Urban Big Data, Multimodal Fusion and Spatial Intelligence Analysis (Second Author).
Forging a Strong Sense of Community for the Chinese Nation Exhibition Project (Participant).
Academic Participation
Reviewer
Conference workshops
- ICLR · LMRL2026
- ICLR · FM4Science2026
- ICML · AI4Science2026
- NeurIPS · MATH-AI2026 · 2025
- NeurIPS · AI for Science2026 · 2025
- ICML · AI for Math2025
- ICLR · AI for Nucleic Acids2025
- ICLR · GenAI Watermarking2025
Learning Experience- 2026.04
- 2025.07
Tsinghua University–Peking University Center for Life Sciences (Tsinghua)
- 2025.07–2025.09
Shenzhen Medical Academy of Research and Translation / Shenzhen Bay Laboratory
- 2025.01
CFPU / Brown University Department of Physics
- 2024.09–2024.12
Chinese Association for Artificial Intelligence (CAAI)
- 2024.08
- 2024.08
- 2024.06–2027.06
Tsinghua University Qian Xuesen Mechanics Class · Shenzhen X-Institute
- 2024.03–2024.06
National Tianyuan Mathematics Central Center, Wuhan University
Internship Experience
Beijing Frontier Research Center for Biological Structure, Tsinghua University 2025.06-present: I developed the Polypeptide Structure Database and have been engaged in work on protein and polypeptide design. Supervised by Prof. Yafei Yuan and Prof. Tong Wang.
Huawei Mindspore Community & Chinese Academy of Sciences Institute of Software Open Source Internship 2024.09-2025.3: I utilized machine learning, artificial intelligence, and other technologies to implement VGG19-based Pollock style transfer paintings with fractal and turbulence feature extraction and NFT tag generation. Currently reported by Huawei Official Account
Bose Quantum "Spark People" Community Intern 2024.09-present
Competition Honors
China Undergraduate Life Science Contest: National Second Prize, 2026
“Challenge Cup” China College Students' Extracurricular Academic Science and Technology Works Competition: Autonomous Region Second Prize, 2026
China College Students' E-Commerce “Innovation, Creativity and Entrepreneurship” Challenge: Autonomous Region Second Prize, 2026
UN Global Leadership and ESG Programme × X-Institute Shenzhen "AI for the Young and the Elderly" Innovation Challenge: Best Innovation Award, 2026
Contemporary Undergraduate Mathematical Contest in Modeling (CUMCM) National Second Prize, 2025.11
Smart Chemical City Challenge (Xinjiang Station): Second Prize, 2025
SynBio Challenges Silver Award, 2025 See Link
The Mathematical Contest in Modeling (MCM) Honorable Mention, 2025.5
2024 Alibaba Cloud Tianchi University Student Competition National Finals 17th Place, 2024.10See Link
2024 14th APMCM Asia-Pacific Mathematical Modeling Competition National Third Prize, 2024.08
"Alpha Egg Cup" 2024 National Amateur Go King Championship and "Commercial Travel Grand Canal Cup" City Go Competition 15th Place, 2024.07
Xinjiang Youth Amateur Go Dan Level Competition 53rd Place, 2024.05
Hunan Province Spring Cup Go Competition 7th Place, 2024.02
National Youth Intellectual Sports Meeting Go Competition 9th Place, 2024.02
Xinjiang University Vulnerability Reporting Honor, 2023.10
2023 Xinjiang "Tianshan Fixed Network Cup" Network Security Skills Competition 7th Place, 2023.10
Website Development
Peptide Design Competition, Tsinghua University: https://www.fbs.frcbs.tsinghua.edu.cn/competition/2025Peptide/
Polypeptide Structure Database: https://www.frcbs.tsinghua.edu.cn/cpdb/
Tong Wang Research Group: https://wanggroup.ai/
About TEEP - Shenzhen X-Institute
Shenzhen X-Institute originated from the Tsinghua University "Xuetang Plan" Qian Xuesen Mechanics Class (abbreviated as "Tsinghua Qian Class"). Founded in 2009, Tsinghua Qian Class is one of the 66 pilot projects of the National "Outstanding Student Training Experimental Program for Basic Disciplines" and the "Tsinghua Xuetang Talent Training Plan". It is the only experimental class not positioned in a single discipline, but focused on engineering foundations (or interdisciplinary innovation across mechanics and all engineering technology disciplines). Its mission is to discover and cultivate innovative talents who aspire to change the world and benefit humanity through technology, explore future innovative talent training models, and answer "Qian Xuesen's Question".
Curriculum Vitae Download.
|