Repository navigation
Tutorial Processing ZTScan data
SCIEX's ZTScan DIA is a fascinating acquisition mode implemented on the 8600 and the 7600+ ZenoTOF systems.
Classical DIA/SWATH approaches use a set of overlapping Q1 isolation windows of 10–50 Da (depending on the vendor, instrument, speed, ...) to fragment virtually every MS1 feature in chunks. Compared to DDA, which uses Q1 isolation windows of ca. 1–2 Da, DIA fragments larger packets of precursors simultaneously and measures the combined MS2 spectrum that includes fragments from all precursors. The wider the Q1 window, the more complex the MS2 spectrum.
DIA elegantly bypasses some of the fundamental limitations of DDA workflows (limited coverage, dead-time to process MS1 scans and decide what to fragment, avoiding isotopes and adducts, triggering at the peak apex, avoiding background, ...), but comes with new challenges in data processing. MS2 must be extracted to support the identification of features (i.e. of the peaks discovered in the MS1×RT space), but need to be cleaned up to remove MS2 fragments that originate from neighboring precursors with similar m/z and RT. Because of the nature of the fragments, this is a step in which proteomics and metabolomics must adopt different strategies.
One strategy is to quantify the correlation over time between the precursor intensity (the EIC of MS1 scans) and the profiles of the fragments from MS2. This works with peptides, but poses an important risk in metabolomics and lipidomics. EIC correlation filtering tends to remove diagnostic fragments that are common to a class, e.g. the 184.07 fragment for PCs, which elute as a continuum giving rise to a plateau of 184 fragments. Applying too stringent a cutoff on the EIC correlation removes a key diagnostic fragment and leads to misannotation.
Additional criteria are necessary alongside EIC correlation. Besides ion mobility (which we don't support), ZTScan offers an orthogonal way of filtering fragments based on MS2 data. Instead of hopping between windows in discrete jumps, in ZTScan the isolation window moves continuously across the mass range of interest, collecting MS2 spectra with small steps in the Q1 isolation window. This generates a triangular envelope in the intensity of fragments, with a maximum aligned with the m/z of the precursor, i.e. when its transmission through Q1 into the collision cell is optimal. This "triangular" envelope across neighboring MS2 spectra allows assessing the alignment of fragments with the precursor.
Calculating the correlation with the precursor and the "Q1 triangularity" for one fragment requires processing a large number of neighboring scans, both MS1 and MS2: slicing spectra, aligning, correlating, ... Considering that a single sample might include 10,000 features, and for each feature it's common to have 10–100 fragments, the operation may have to be repeated up to 1 million times per sample. It can be very costly, both in terms of computational time and memory requirements.
Since none of the commonly used metabolomics software was able to swiftly process ZTScan data (see hours of waiting time, crashing, memory issues, ...), we set out to provide a faster, scalable, and nevertheless quantitative alternative.
MASSter provides several special functionalities to analyze ZTScan data:
- We embedded quantitative metrics to score how each individual MS2 fragment is correlated to the precursor intensity in MS1 scans (
eic_corr) or with the Q1-isolation window across MS2 scans (q1_ratio, a measure of "triangularity"). These are calculated during processing and stored. When exporting MS2 spectra withexport_mgf(), the user can decide whether and how to filter MS2 fragments. - To calculate
eic_corrandq1_ratio, MASSter automatically pulls profile MS1 and MS2 scan data, calculates centroids, aligns them, and estimates the metrics. Despite typically having hundreds of thousands of MS2 fragments to score, we optimized the procedure to take less than 1 minute (normally) for each file. - Beyond speed, we also optimized for low memory usage. ZTScan files are much larger and include hundreds of thousands of MS2 spectra (up to 1 million). Nevertheless, the memory demand is comparable to that of a DDA file. This allows analyzing multiple samples in parallel and running the Wizard without special adjustments.
- MASSter works directly on raw data: wiff or wiff2. There is no need to convert to large mzML files before processing.
- Beyond mere processing, we also provide some special plots tailored to ZTScan data.
Disclaimer: the analysis of ZTScan (and DIA in general) data is a matter of active research. We continuously test, optimize, and hopefully improve the results. It's possible that outcomes will change across MASSter versions and, therefore, you might want to pin the environment. If you use the Wizard (recommended), a pyproject.toml will be automatically created to control the environment.
We first describe the functionalities for a single file. Note: you might want to first read how MASSter analyzes single files in general. It's not a precondition, but it helps in understanding the steps. Later, we explain how to achieve the same in the Wizard.
The only important thing you have to do is add type='ztscan' when creating the Sample. That's it!
We illustrate this with an example *.wiff file (Note: the file is way too large to be attached to this repository). A summary script is at the end.
A file with >100,000 MS scans loads in under 20 seconds (!). During this period, all MS1 scans have been centroided.

Next, we detect features, adducts, and isotopic patterns. We then link MS2 spectra to features. All these operations are executed with default settings. MASSter recognizes it's a ZTScan file and will automatically calculate eic_corr and q1_ratio for all MS2 fragments of 1141 (in yellow). The total number of features is 1146, but 5 are outside of the m/z range that was considered for fragmentation.
All these steps take 1.5 minutes, almost entirely for the last step. As shown in the test, it takes about 0.1 s per feature (this time includes identifying the neighboring spectra, slicing, centroiding, aligning, and calculating the two metrics). The total time scales with the number of features to analyze. For a sample with 10,000 features and 500,000+ MS2 fragments, the time can increase to about 10 minutes. These computations are done on a single core/thread, with a negligible extra cost in memory compared to a DDA file.

At this point, virtually all features have an associated MS2 spectrum including eic_corr and q1_ratio scores. No filter has yet been applied. Features can be visualized...

... and their details are available in the features_df dataframe. In a notebook, this can be used to quickly sort and filter rows to explore the data:

The MS2 spectra are stored in the column ms2_spec as Spectrum object. They can be obtained for any feature using the feature_id or the feature_uid and the get_ms2() method. The result is a Spectrum object, which can be exported to a dictionary with .to_dict() or visualized with .plot():

The eic_corr and q1_ratio are also present and are visible in the hover tooltip. For development, we integrated the option to color the spectrum by any property, including eic_corr or q1_ratio:

For testing, MASSter also offers the option to plot the chromatographic profiles of the precursors and all MS2 fragments (which are used to calculate the eic_corr), or the intensity across the ZTScan windows (for the q1_ratio):

Most importantly, MS2 spectra for all fragments can be exported with export_mgf(). At this step, it's possible to apply thresholds to export only a subset of MS2 peaks. This allows, for example, testing different cutoffs (eic_corr_min and q1_ratio_min) and creating multiple files for testing, using filename= to adjust the name if necessary. MS2 spectra can be annotated with the tool of interest. In our case, we annotated the features with tima and reimported the results for visualization. MASSter also includes MS2-matching in case a library is available as MGF or MSP.

This illustrates the basic function for ZTScan analysis. In brief, the only key steps are to (i) use type='ztscan' when instantiating the Sample and (ii) add cutoffs when exporting with export_mgf(). As shown in this example, processing a *.wiff file to obtain a list of features and filtered MS2 spectra is a matter of minutes on a single processor.
This is a summary of the steps typically used to analyze a single ZTScan file:
from masster import Sample
sample = Sample(filename='....wiff', type='ztscan')
# feature detection
sample.find_features(noise=100, chrom_fwhm=1.0)
# optional: remove features with poor chromatographic shape
sample.features_filter(coherence=0.4, prominence=2.0)
# optional: detect adducts and obtain isotopic data
sample.find_adducts()
sample.find_iso()
# essential: MS2 link
sample.find_ms2()
# save everything to sample5
sample.save()
# export
sample.export_mgf("ztscan_filtered.mgf", q1_ratio_min=0.3, eic_corr_min=0.3) # change cutoffs as needed
sample.export_excel('mydata.xlsx')
sample.export_mztab('mydata.mztab')We recommend using the Wizard as one would do for DDA, and apply two small changes:
from masster import Wizard
wiz = Wizard(source="./folder_with_wiff",
folder="./output_folder",
num_cores=8)
wiz.create_scripts() # analyze one sample and create *.py scriptsThen, edit 1_processing.py introducing the same modifications described above, namely
- (line ~179): Replace
sample.load(filename=str(raw_file))withsample.load(filename=str(raw_file), type='ztscan')1b. (line ~158): Replacesample.load(filename=str(raw_file), sample_idx=sample_idx)withsample.load(filename=str(raw_file), sample_idx=sample_idx, type='ztscan')(this is used to deal with multiple injections in the samewifffile) - (line ~344): Replace
study.export_mgf()withstudy.export_mgf(q1_ratio_min=0.3, eic_corr_min=0.3)
Then run the Wizard with:
from masster import Wizard
wiz = Wizard(source="./folder_with_wiff",
folder="./output_folder",
num_cores=8)
wiz.test_and_run() # analyze one sample and create *.py scriptsor simply
uv run python "...../1_notebook.py" # update path
And then: relax and wait...
As it is simple to parallelize the time-consuming process of calculating eic_corr and q1_ratio, it's easy to process large studies with numerous ZTScan files in a matter of hours with limited memory requirements.
As reference, a recent test with 9 ZTScan files with >600,000 MS scans each, 23 GB in total, completed the analysis in about 15 minutes with 10 cores. Overall, the analysis time is marginally longer than with DDA!
