AI Proteomics Initiative
AI-Ready Mass Spectrometry (AIR-MS)
Unlocking Hidden Proteomics Data: Turning In-house MS Data into AI-ready Resources
AIR-MS aims to build a large-scale, diverse, and standardized AI-ready proteomics resource by transforming valuable proteomics datasets generated worldwide into reusable resources for AI model development and benchmarking.
We have established the first dataset, MassNet, an AI-ready proteomics resource comprising ∼28,000 DDA-MS files and ∼46,000 DIA-MS DIA-MS files. The MassNet [1] and its associated work have been accepted in principle by Nature Methods.
We welcome contributions from proteomics laboratories worldwide.
Participate / 参与计划
AIR-MS Registration Form
AIR-MS计划参与登记表
Submit your lab and project information. The lab code and password will be issued after review.
填写实验室与项目信息,审核通过后获得 lab code 与密码。
Fill the form / 前往登记Data Sharing Link Submission
数据分享地址提交
Log in with your lab code and password, then submit the dataset sharing link and track its status.
使用 lab code 与密码登录后提交数据集下载地址,并可查看下载状态。
Go to submission / 前往提交Project Progress Board
项目参与进度看板
Browse the contribution of all participating labs: lab code, country/region and files collected.
查看所有参与实验室的贡献情况:lab code、国家地区与已收集文件数。
View progress / 查看进度1. Dataset Requirements
We currently prioritize global proteomics datasets. Specialized datasets, including PTM proteomics, immunopeptidomics, and metaproteomics, are also welcome, with additional metadata required for downstream processing and database searching.
| Item | Requirements |
|---|---|
| MS Acquisition | Label-free DDA and DIA |
| Species | No restriction |
| Sample Type | Biological research samples |
| Mass spectrometers |
|
2. Required Metadata
Please provide the following experimental information for each dataset. Additional details may be required for specialized datasets to support appropriate FASTA/database selection and downstream database searching.
| Raw file name | e.g., sample_001.raw |
| Dataset description | e.g., Human plasma proteomics dataset |
| Species | e.g., Homo sapiens |
| Biological condition | e.g., Benign tissue |
| Instrument | e.g., Orbitrap Exploris 480 |
| Acquisition mode | e.g., DDA |
| Digestion | e.g., Trypsin |
| Additional information (if applicable) | e.g., PTM type / HLA-I or HLA-II / metaproteomics sample source |
Metadata template: Download here
3. Excluded Data
The following datasets are not required:
| ❌ QC samples |
| ❌ System suitability test samples |
| ❌ Instrument performance monitoring runs |
4. Data Submission
1. Prepare Data
- Raw MS files
- Corresponding metadata information
2. Upload Data
Upload raw files to your preferred secure storage platform (e.g., Google Drive, Baidu disk) and share the download link with us.
3. Data Processing
Our team will perform: metadata harmonization, standardized processing, quality control, annotation, and conversion into AI-ready formats.