Skip to content

Repository files navigation

Hanns

High-performance approximate nearest neighbor (ANN) search in pure Rust.

Hanns is the ANN engine powering three production systems:

Ecosystem Role
Milvus Drop-in replacement for KnowWhere C++ in the world's most popular cloud-native vector database
HannsDB Embedded ANN engine for single-machine agent workloads — low-latency, zero-dependency
Lance ANN backend for the open multimodal vector lake format

Built from scratch. No C++ dependencies. Benchmarked head-to-head against FAISS, KnowWhere, and Lance on real x86 server hardware.


Performance at a Glance

Fresh HannsDB-x86 vs official Zilliz Knowhere (SIFT-1M, top_k=100)

The table below summarizes the fresh authority runs against official zilliztech/knowhere. All rows were run on the remote HannsDB-x86 authority machine and validated from archived logs/status files under docs/parity/.

Claim boundary: these are scoped verdicts, not a blanket "Hanns beats Knowhere everywhere" claim. In particular, DiskANN/AISAQ remains native_comparable=false, so it is reported as a constrained numeric observation rather than a family leadership verdict.

Family / lane Hanns result Official target Conclusion Evidence
HNSW near-0.80 R@100 0.9198, QPS 30,972, build 23.88s R@100 0.9178, VPS 30,014, build 29.34s Scoped win docs/parity/hanns-knowhere-hnsw-verdict-20260424T162100Z.json
HNSW near-0.95 R@100 0.9506, QPS 27,220, build 28.88s R@100 0.9500, VPS 23,370, build 29.34s Scoped win docs/parity/hanns-knowhere-hnsw-near095-verdict-20260425T021000Z.json
HNSW-SQ emitted FP32 row R@100 0.9595, QPS 38,562, build 59.49s R@100 0.9531, VPS 30,960, build 66.28s Scoped win docs/parity/hanns-knowhere-hnsw-sq-verdict-20260425T090100Z.json
HNSW-PQ raw-candidate/raw-refine fair-8 R@100 0.9758, QPS 9,451, build 78.94s R@100 0.9576, VPS 5,051, build 84.54s Custom-variant win; not same-semantics family leadership docs/parity/hanns-knowhere-hnsw-pq-raw-refine-fair8-verdict-20260425T114300Z.json
IVF-PQ R@100 0.9050, QPS 7,520, build 6.43s R@100 0.7841, VPS 1,399, build 8.96s Scoped win docs/parity/hanns-knowhere-ivfpq-buildtime-verdict-20260424T151136Z.json
IVF-SQ8 near-0.80 R@100 0.8104, QPS 39,521, build 3.11s R@100 0.8009, VPS 28,374, build 6.88s Scoped win docs/parity/hanns-knowhere-ivfsq8-near080-verdict-20260425T030700Z.json
IVF-SQ8 near-0.95 R@100 0.9532, QPS 15,690, build 3.11s R@100 0.9519, VPS 13,056, build 6.88s Scoped win docs/parity/hanns-knowhere-ivfsq8-near095-verdict-20260425T030700Z.json
IVF-USQ / RabitQ search R@100 0.6144, QPS 19,662 R@100 0.6146, VPS 7,728 Scoped search-throughput win at matched recall band docs/parity/hanns-knowhere-ivfusq-search-verdict-20260425T064738Z.json
DiskANN/AISAQ page-cache expanded R@100 0.9937, QPS 624.9, build 124.19s, persist 62.22s AISAQ_S near-0.95: R@100 0.9515, VPS 520.39, build 159.80s Numeric observation only; native_comparable=false docs/parity/hanns-knowhere-diskann-aisaq-page-cache-expanded-verdict-20260426T024300Z.json

DiskANN/AISAQ details: the best constrained row uses search_surface=page_cache, disk_pq_dims=32, pq_candidate_expand_pct=400, and rerank_expand_pct=400, with scope_audit.has_page_cache=true. Because the implementation is still a constrained PQFlashIndex skeleton rather than a proven native-comparable SSD DiskANN/AISAQ pipeline, the validator keeps leadership_claim_allowed=false.

vs KnowWhere C++ inside Milvus (Cohere Wikipedia-1M, 768-dim IP, x86)

Metric KnowWhere C++ Hanns vs Native
Graph build (Optimize) 854.2s 336.9s 2.53× faster
Index load 1158.9s 673.7s 1.72× faster
QPS (c=20, ef=128, k=100) ~500 1,051 ~2.1× faster
QPS (c=80, ef=128, k=100) ~800 1,042 ~1.3× faster
Recall@100 0.960 0.957 parity

vs Lance (Cohere-1M, 1024-dim L2, x86)

ef Hanns QPS Lance QPS Hanns/Lance
50 2,331 1,473 1.58×
200 794 483 1.64×
800 245 140 1.75×

Hanns search leads by 1.58–1.75× across the full ef range; advantage grows with higher recall requirements.

vs Microsoft DiskANN Rust (SIFT-1M, R=48 L=64, x86)

System Recall@10 QPS
DiskANN Rust (Microsoft) 0.986 4,832
Hanns AISAQ NoPQ 0.994 5,806

+20% higher QPS, +0.8% better recall at identical parameters on identical hardware.


Advanced Quantization: USQ

Standard Product Quantization (PQ) — used by FAISS and most ANN libraries — minimizes L2 reconstruction error. On modern embedding search (high-dim, Inner Product metric), this is the wrong objective: PQ recall collapses near zero.

USQ (Unit Sphere Quantizer) applies a QR orthogonal rotation before quantizing, making compression metric-agnostic. Derived from RabitQ principles and extended to multi-bit (1/4/8-bit) with:

  • QR decomposition rotation matrix (trained once, applied to all vectors)
  • Unified 1/4/8-bit scalar quantization in the rotated space
  • AVX512VNNI integer dot product scoring (vpdpbusd)
  • Two-stage pipeline: 1-bit FastScan filter → B-bit rerank

USQ Quantization

Dataset: Cohere Wikipedia-1M (768-dim, Inner Product), nprobe=32, x86 authority.

Method Compression QPS Recall@10 Usable?
IVF-PQ m=32 8× 723 0.066 ✗
IVF-PQ m=48 5.3× 502 0.127 ✗
IVF-Flat 1× 339 0.798 ✓
IVF-SQ8 4× 605 0.805 ✓
USQ 4-bit 8× 1,308 0.879 ✓
USQ 8-bit 4× 1,011 0.968 ✓

At the same 8× compression ratio where PQ gives recall=0.066, USQ gives 0.879 — a 13× improvement. USQ 8-bit simultaneously delivers 3× faster QPS than IVF-Flat with +17% better recall at ¼ the memory.

On 3072-dim embeddings (SimpleWiki-OpenAI-260K), USQ 8× still achieves recall 0.925 at 1,607 QPS.


Ecosystem Integration

Hanns ships as three distinct integration surfaces from a single codebase:

crate-type = ["staticlib", "cdylib", "rlib"]
     │               │              │
  Milvus C ABI    Python/JNI    HannsDB / Lance (native Rust)

Milvus — C ABI drop-in

Integration method: src/ffi.rs exposes 31 #[no_mangle] extern "C" functions that mirror the KnowWhere C++ API exactly (knowhere_create_index, knowhere_add_index, knowhere_search, knowhere_search_with_bitset, …). Milvus links against the compiled staticlib/cdylib with zero source changes — same data format, same query semantics, same bitset filtering interface.

Milvus standalone
  └─ KnowWhere shim (C++)
       └─ dlopen libhanns.so   ← same ABI as native KnowWhere
            └─ IndexWrapper::dispatch → HnswIndex / IvfSq8Index / PQFlashIndex / …

The IndexWrapper is the central dispatch struct — a single IndexKind enum that routes all C API calls to the correct Rust index type at zero overhead (no vtable, no allocation on the hot path).

VectorDBBench end-to-end results confirm 2× QPS across HNSW, IVF-Flat, IVF-SQ8, and IVF-PQ under realistic concurrent load:

Round Change QPS (c=80)
R4 FFI lazy bitset allocation 349
R7 Private rayon ThreadPool (HNSW_NQ_POOL) 540
R8 Eliminate BinaryHeap clone + pre-alloc output buffer 1,042

HannsDB — native Rust crate

Integration method: HannsDB adds Hanns as a Cargo dependency and uses index types directly:

use hanns::{HnswIndex, IvfFlatIndex, IvfUsqIndex};

No FFI boundary, no serialization overhead. HannsDB wraps index instances in its own VectorIndexBackend trait and handles persistence, collection management, and the VectorDBBench client layer. VectorDBBench authority results (x86, 1536-dim, k=100):

Metric Result
Load 148.0s
Optimize 87.9s
p99 latency 1.8ms
Recall 0.9756

Previously, cosine search p99 reached 110ms due to per-query allocations. After fixing TLS scratch buffer reuse: p99 = 3.5ms (31× improvement).

Lance — native Rust trait impl

Integration method: Lance defines an IvfSubIndex trait for pluggable ANN backends. Hanns implements it for HnswIndex, letting Lance's vector lake query engine call Hanns search directly in-process:

// in Lance repo
impl IvfSubIndex for hanns::HnswIndex { … }

No serialization boundary, no IPC. Geometry-mean speedup across ef=50–800: 1.64× with equivalent recall. Build is ~26% slower (Lance uses rayon parallel build; Hanns parallel build improvement is in progress).

Lance also wires Hanns-backed IVF_USQ behind cfg(feature = "hanns") in the vector builder / IVF stack, which is what powers the 1B large-top (k=10,000) runs below.

1B large-top capability (PCA-512, cosine, 5 × 194M shards, x86 ECS):

Tier Hanns-backed Lance config Recall@10K Mean latency What it demonstrates
DRAM IVF_USQ 4-bit np=256 rf=2 0.934 957ms Best sub-0.95 local-storage point; slightly faster than Lance IVF_RQ at the same recall band
DRAM IVF_USQ 8-bit np=256 rf=1 0.973 1,826ms Practical high-recall point, but the 8-bit auxiliary index is too large to be a clear DRAM default
SSD IVF_USQ 4-bit np=256 rf=2 0.934 964ms Near-DRAM latency on cold NVMe
OBS IVF_USQ 4-bit np=256 rf=1 0.803 6,555ms Lowest-latency remote-storage frontier point
OBS IVF_USQ 8-bit np=256 rf=1 0.973 7,710ms Collapses the old 0.81-0.96 OBS gap; faster than Lance IVF_SQ (10,518ms) and much faster than Lance IVF_RQ rf=2 (19,306ms) while also improving recall
OBS IVF_USQ 8-bit np=1024 rf=1 0.9949 21,312ms Best sub-0.995 remote point

At the extreme high-recall end, Lance + Hanns reaches 0.9967 recall@10K on OBS at np=1024 rf=2 in 27,988ms. Build is already production-scale as well: the 5-shard IVF_USQ 8-bit index finishes in about 57 min with 117GB auxiliary.idx per shard (about 1.3TB peak temp space).

The practical reading of these numbers is simple:

  • On OBS / remote object storage, Hanns-backed IVF_USQ is the key Lance large-top enabler.
  • On DRAM / SSD, 4-bit USQ gives Lance strong sub-0.95 points, while 8-bit USQ trades much larger index size for higher recall.

Python / JVM bindings

Python (feature = "python"): PyO3 bindings in src/python/ expose HnswIndex, IvfFlatIndex, IvfPqIndex, IvfUsqIndex, MemIndex as native Python classes.

JVM / Android (feature = "jni-bindings"): JNI bindings in src/jni/ expose the same index set via @NativeMethod for Java/Kotlin callers.


Performance Charts

QPS Comparison

Recall vs QPS

DiskANN Comparison

All benchmarks run on dedicated x86 server, target-cpu=native. Apple Silicon builds are for fast iteration only — not used as authority evidence.


Index Coverage

Index Status Recall@10 Notes
HNSW ✅ Leading 0.972 (SIFT-1M) +11.9% vs FAISS 8T; 2.53× faster build in Milvus
HNSW-SQ ✅ Ready 0.992 Integer precomputed ADC path
IVF-Flat ✅ Leading 0.978 (SIFT-1M) 5.2× faster than FAISS 8T (batch parallel)
IVF-SQ8 ✅ Leading 0.958 (SIFT-1M) 1.42× faster than FAISS 8T; AVX2 fused decode
IVF-USQ ✅ Ready 0.905–0.968 (Cohere-1M) AVX512VNNI; unified 1/4/8-bit
IVF-PQ ✅ Ready varies m=32: 0.720 on synthetic data
AISAQ (DiskANN Flash) ✅ Ready 0.994 NoPQ (SIFT-1M) On-demand pread + io_uring; 6.5× faster build than native
ScaNN ✅ Ready 0.969 Exceeds 0.95 gate at reorder_k=1600
Sparse / WAND ✅ Ready 1.0 Sparse vector retrieval
Binary ✅ Ready — Hamming distance

Quantization Subsystem

src/quantization/
  usq/           UsqQuantizer — QR rotation + unified 1/4/8-bit quantization
    rotator.rs   QR decomposition rotation matrix
    quantizer.rs training + SIMD scoring (AVX512VNNI)
    fastscan.rs  AVX512 fast scan (1-bit stage) + topk
    searcher.rs  two-stage coarse filter + rerank
  pq/            Product Quantizer — parallel k-means
  sq/            Scalar Quantizer — SQ8/SQ4
  pca/           PCA transform — nalgebra SVD (pure Rust, no BLAS)

Build

cargo build --release          # LTO + codegen-units=1 + target-cpu=native
cargo test
cargo run --example benchmark --release

.cargo/config.toml enables target-cpu=native on x86_64 and aarch64 automatically.


Repository Layout

src/
  faiss/           core index implementations
  quantization/    quantization subsystem (USQ, PQ, SQ, PCA)
  ffi/             FFI layer (KnowWhere-compatible ABI)
tests/             integration and regression tests
examples/          full benchmark harness
assets/benchmarks/ comparison charts
docs/              design docs, performance audits
benchmark_results/ authority verdict artifacts (JSON)
scripts/           chart generation, remote build/test
wiki/              operational runbooks and authority numbers

Datasets

Dataset Dim Metric Size
SIFT-1M 128 L2 1M vectors
Cohere Wikipedia-1M 768 IP 1M vectors
Cohere-1M 1024 L2 1M vectors
SimpleWiki-OpenAI-260K 3072 IP 260K vectors

Authority Hardware

Performance numbers are produced on a dedicated x86 server with target-cpu=native (specs). Apple Silicon builds are for fast iteration and pre-screening only.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages