Repository navigation
Tutorial Identification
Zamboni Nicola edited this page Jan 29, 2026
·
2 revisions
This tutorial describes identification workflows around MASSter outputs.
Three common patterns:
- MS1 library matching inside MASSter (fast, broad, lower confidence).
- MS1+MS2 identification with identify2() (integrated spectral matching).
- MS2 identification in external tools (higher confidence, tool-specific).
Typical flow in a Study:
# Load a built-in library (examples: "hsapiens", "scerevisiae", "ecoli")
study.lib_load("hsapiens")
# Identify consensus features via MS1 matching
study.identify()Results are written to identification tables and also summarized into "top ID" columns where applicable.
MASSter includes built-in identify2() for MS1/MS2 spectral matching using simsimd.
Requirements:
- Install simsimd:
pip install simsimd - Load a library with MS2 spectra (e.g., via
import_mgf())
Basic usage:
# Load library with MS2 spectra
study.import_mgf("spectral_library.mgf")
# MS2-only identification (default)
study.identify2()
# Combined MS1 and MS2 identification
study.identify2(mslevel=[1, 2])
# MS1-only identification
study.identify2(mslevel=[1], ms1_target='all')Key parameters:
-
mslevel: List of MS levels[1, 2](default:[2]) -
min_score: Minimum cosine similarity (default: 0.5) -
mz_tol_ms2: m/z tolerance for MS2 in Da (default: 0.02) -
mz_tol_ms1: m/z tolerance for MS1 in Da (default: 0.01) -
metric: Spectral similarity -'cosine','modified-cosine', or'combo'
Example with custom parameters:
study.identify2(
mslevel=[1, 2],
min_score=0.7, # Higher stringency
mz_tol_ms2=0.01, # Tighter MS2 tolerance
metric="combo", # Use both cosine and neutral loss
only_orphans=True # Only identify unmatched features
)MASSter's role here is to export the right files:
-
consensus.csv(feature table for metadata and filtering) -
consensus.mgf(spectra for database search / annotation)
After running the external tool, you typically:
- copy the tool output into your project,
- import results back into MASSter,
- filter/refine annotations,
- update exported tables/plots.
- Export MGF + consensus table (Wizard pipelines usually do this already).
- Run TIMA outside MASSter.
- Import results:
study.import_tima("path/to/tima_output")- Export MGF.
- Run LipidOracle outside MASSter.
- Import results:
study.import_oracle("path/to/oracle_output")Common operations:
# Full joined identification table (id_df + library columns)
id_df = study.get_id()
# Select and filter identifications (examples)
high_conf = study.id_select(score=(0.8, 1.0))
study.id_filter(high_conf)
# Recompute/refresh "top ID" summary columns in consensus/features tables
study.id_update()To understand the exported artifacts, see: