Skip to content

Latest commit

 

History

History
42 lines (27 loc) · 3.28 KB

File metadata and controls

42 lines (27 loc) · 3.28 KB

Dependency Audit Toolchain Usage Guide

This document outlines operational procedures for the Dependency Audit Crawler and the Dependency Dashboard Visualizer.

Architecture Overview

The toolchain consists of two discrete components:

  1. The Crawler (Backend): A Python-based script that identifies downstream dependencies via Sourcegraph and extracts metadata using the GitHub GraphQL API.
  2. The Dashboard (Frontend): A static HTML and JavaScript application that parses crawler output into readable tables, network graphs, and citation lists.

The Crawler

The crawler searches global open-source indices to find projects utilizing a target repository. It outputs a Universal Dependency Graph (.json) and a set of SPDX 2.3 manifests.

For each project it compiles an identifier set (header, CMake/Bazel/pkg-config, and repository-URL identifiers) from the project's own files and searches for the ways consumers reference them. Every dependency edge in the graph carries a confidence tier and score, the evidence and identifiers behind it, its provenance, and a relationship label (DEPENDS_ON, or VENDORED/MIRROR for bundled copies and forks) — so the dashboard can sort and filter dependents by how strongly the evidence supports them. See OVERVIEW.md for the methodology.

Data Generation

The crawler is deployed primarily as a GitHub Action. Upon execution, the action outputs an archive containing all graph data and SPDX manifests. This archive is required to initialize the dashboard.

  1. Navigate to the target repository's GitHub Actions panel.
  2. Select the Dependency Audit Crawler workflow.
  3. Trigger the workflow manually, supplying the required project_name parameter.
  4. Download the resulting artifact archive once the workflow completes.

The archive contains a flat file structure: the dependency_graph.json mapping file, and an spdx_snippets directory containing strictly version-pinned SBOMs.

The Dashboard Visualizer

The visualizer consumes the archive generated by the crawler and processes it strictly client-side. No data is transmitted externally during viewing.

Initialization

  1. Open the dashboard.
  2. Select Upload Data and provide the .zip archive downloaded from the GitHub Action.
  3. The application will initialize an extraction process and load the data into memory.

Interface Components

  • Dependents List: The default view. It displays a paginated data table of all identified consumers. Filtering options include grouping by organizational owner, isolating leaf nodes, sorting by metadata parameters, and filtering by maximum chain depth.
  • Network Graph: An interactive mapping of the dependency tree. The interface supports orthogonal routing for dense graphs and radial displays for heatmaps. Search functionality and node-count limits restrict rendering to preserve performance on large datasets.
  • Citations: An aggregated view of DOI links and academic papers extracted from downstream dependents.
  • Inspector Panel: Persists on the right side of the interface. Selecting any node or list item populates this panel with deeper repository statistics and renders the associated SPDX manifest for the dependency relationship. Single manifests or the complete bulk archive can be downloaded directly from this panel.