AI-Ready Mass Spectrometry (AIR-MS)
Unlocking Hidden Proteomics Data: Turning In-house MS Data into AI-ready Resources
MassNet aims to build a large-scale, diverse, and standardized AI-ready proteomics resource by transforming valuable proteomics datasets generated worldwide into reusable resources for AI model development and benchmarking.
We have established the first version of MassNet, an AI-ready proteomics resource comprising ∼28,000 DDA-MS files and ∼46,000 DIA-MS files. The MassNet-DDA [1] and its associated work have been accepted in principle by Nature Methods.
We welcome contributions from proteomics laboratories worldwide.
1. Dataset Requirements
We currently prioritize global proteomics datasets. Specialized datasets, including PTM proteomics, immunopeptidomics, and metaproteomics, are also welcome, with additional metadata required for downstream processing and database searching.
| Item | Requirements |
|---|---|
| MS Acquisition | Label-free DDA and DIA |
| Species | No restriction |
| Sample Type | Biological research samples |
| Mass spectrometers |
|
2. Required Metadata
Please provide the following experimental information for each dataset. Additional details may be required for specialized datasets to support appropriate FASTA/database selection and downstream database searching.
| Raw file name | e.g., sample_001.raw |
| Dataset description | e.g., Human plasma proteomics dataset |
| Species | e.g., Homo sapiens |
| Biological condition | e.g., Benign tissue |
| Instrument | e.g., Orbitrap Exploris 480 |
| Acquisition mode | e.g., DDA |
| Digestion | e.g., Trypsin |
| Additional information (if applicable) | e.g., PTM type / HLA-I or HLA-II / metaproteomics sample source |
Metadata template: Download here
3. Excluded Data
The following datasets are not required:
| ❌ QC samples |
| ❌ System suitability test samples |
| ❌ Instrument performance monitoring runs |
4. Data Submission
1. Prepare Data
- Raw MS files
- Corresponding metadata information
2. Upload Data
Upload raw files to your preferred secure storage platform (e.g., Google Drive, Baidu disk) and share the download link with us.
3. Data Processing
Our team will perform: metadata harmonization, standardized processing, quality control, annotation, and conversion into AI-ready formats.