Skip to content

Tutorial Identification

Zamboni Nicola edited this page Jan 29, 2026 · 2 revisions

Tutorial: Identification

This tutorial describes identification workflows around MASSter outputs.

Three common patterns:

  1. MS1 library matching inside MASSter (fast, broad, lower confidence).
  2. MS1+MS2 identification with identify2() (integrated spectral matching).
  3. MS2 identification in external tools (higher confidence, tool-specific).

A) MS1 identification inside MASSter (lib_load + identify)

Typical flow in a Study:

# Load a built-in library (examples: "hsapiens", "scerevisiae", "ecoli")
study.lib_load("hsapiens")

# Identify consensus features via MS1 matching
study.identify()

Results are written to identification tables and also summarized into "top ID" columns where applicable.

B) MS1+MS2 identification (identify2)

MASSter includes built-in identify2() for MS1/MS2 spectral matching using simsimd.

Requirements:

  • Install simsimd: pip install simsimd
  • Load a library with MS2 spectra (e.g., via import_mgf())

Basic usage:

# Load library with MS2 spectra
study.import_mgf("spectral_library.mgf")

# MS2-only identification (default)
study.identify2()

# Combined MS1 and MS2 identification
study.identify2(mslevel=[1, 2])

# MS1-only identification
study.identify2(mslevel=[1], ms1_target='all')

Key parameters:

  • mslevel: List of MS levels [1, 2] (default: [2])
  • min_score: Minimum cosine similarity (default: 0.5)
  • mz_tol_ms2: m/z tolerance for MS2 in Da (default: 0.02)
  • mz_tol_ms1: m/z tolerance for MS1 in Da (default: 0.01)
  • metric: Spectral similarity - 'cosine', 'modified-cosine', or 'combo'

Example with custom parameters:

study.identify2(
    mslevel=[1, 2],
    min_score=0.7,         # Higher stringency
    mz_tol_ms2=0.01,       # Tighter MS2 tolerance
    metric="combo",        # Use both cosine and neutral loss
    only_orphans=True      # Only identify unmatched features
)

C) External MS2 tools (TIMA / LipidOracle)

MASSter's role here is to export the right files:

  • consensus.csv (feature table for metadata and filtering)
  • consensus.mgf (spectra for database search / annotation)

After running the external tool, you typically:

  • copy the tool output into your project,
  • import results back into MASSter,
  • filter/refine annotations,
  • update exported tables/plots.

C1) TIMA (general metabolomics MS/MS)

  1. Export MGF + consensus table (Wizard pipelines usually do this already).
  2. Run TIMA outside MASSter.
  3. Import results:
study.import_tima("path/to/tima_output")

C2) LipidOracle (lipid MS/MS)

  1. Export MGF.
  2. Run LipidOracle outside MASSter.
  3. Import results:
study.import_oracle("path/to/oracle_output")

D) Managing identifications inside MASSter

Common operations:

# Full joined identification table (id_df + library columns)
id_df = study.get_id()

# Select and filter identifications (examples)
high_conf = study.id_select(score=(0.8, 1.0))
study.id_filter(high_conf)

# Recompute/refresh "top ID" summary columns in consensus/features tables
study.id_update()

To understand the exported artifacts, see:

Clone this wiki locally