High-performance approximate nearest neighbor (ANN) search in pure Rust.
Hanns is the ANN engine powering three production systems:
| Ecosystem | Role |
|---|---|
| Milvus | Drop-in replacement for KnowWhere C++ in the world's most popular cloud-native vector database |
| HannsDB | Embedded ANN engine for single-machine agent workloads — low-latency, zero-dependency |
| Lance | ANN backend for the open multimodal vector lake format |
Built from scratch. No C++ dependencies. Benchmarked head-to-head against FAISS, KnowWhere, and Lance on real x86 server hardware.
The table below summarizes the fresh authority runs against official
zilliztech/knowhere. All rows were
run on the remote HannsDB-x86 authority machine and validated from archived
logs/status files under docs/parity/.
Claim boundary: these are scoped verdicts, not a blanket "Hanns beats Knowhere everywhere" claim. In particular, DiskANN/AISAQ remains
native_comparable=false, so it is reported as a constrained numeric observation rather than a family leadership verdict.
| Family / lane | Hanns result | Official target | Conclusion | Evidence |
|---|---|---|---|---|
| HNSW near-0.80 | R@100 0.9198, QPS 30,972, build 23.88s |
R@100 0.9178, VPS 30,014, build 29.34s |
Scoped win | docs/parity/hanns-knowhere-hnsw-verdict-20260424T162100Z.json |
| HNSW near-0.95 | R@100 0.9506, QPS 27,220, build 28.88s |
R@100 0.9500, VPS 23,370, build 29.34s |
Scoped win | docs/parity/hanns-knowhere-hnsw-near095-verdict-20260425T021000Z.json |
| HNSW-SQ emitted FP32 row | R@100 0.9595, QPS 38,562, build 59.49s |
R@100 0.9531, VPS 30,960, build 66.28s |
Scoped win | docs/parity/hanns-knowhere-hnsw-sq-verdict-20260425T090100Z.json |
| HNSW-PQ raw-candidate/raw-refine fair-8 | R@100 0.9758, QPS 9,451, build 78.94s |
R@100 0.9576, VPS 5,051, build 84.54s |
Custom-variant win; not same-semantics family leadership | docs/parity/hanns-knowhere-hnsw-pq-raw-refine-fair8-verdict-20260425T114300Z.json |
| IVF-PQ | R@100 0.9050, QPS 7,520, build 6.43s |
R@100 0.7841, VPS 1,399, build 8.96s |
Scoped win | docs/parity/hanns-knowhere-ivfpq-buildtime-verdict-20260424T151136Z.json |
| IVF-SQ8 near-0.80 | R@100 0.8104, QPS 39,521, build 3.11s |
R@100 0.8009, VPS 28,374, build 6.88s |
Scoped win | docs/parity/hanns-knowhere-ivfsq8-near080-verdict-20260425T030700Z.json |
| IVF-SQ8 near-0.95 | R@100 0.9532, QPS 15,690, build 3.11s |
R@100 0.9519, VPS 13,056, build 6.88s |
Scoped win | docs/parity/hanns-knowhere-ivfsq8-near095-verdict-20260425T030700Z.json |
| IVF-USQ / RabitQ search | R@100 0.6144, QPS 19,662 |
R@100 0.6146, VPS 7,728 |
Scoped search-throughput win at matched recall band | docs/parity/hanns-knowhere-ivfusq-search-verdict-20260425T064738Z.json |
| DiskANN/AISAQ page-cache expanded | R@100 0.9937, QPS 624.9, build 124.19s, persist 62.22s |
AISAQ_S near-0.95: R@100 0.9515, VPS 520.39, build 159.80s |
Numeric observation only; native_comparable=false |
docs/parity/hanns-knowhere-diskann-aisaq-page-cache-expanded-verdict-20260426T024300Z.json |
DiskANN/AISAQ details: the best constrained row uses search_surface=page_cache,
disk_pq_dims=32, pq_candidate_expand_pct=400, and
rerank_expand_pct=400, with scope_audit.has_page_cache=true. Because the
implementation is still a constrained PQFlashIndex skeleton rather than a
proven native-comparable SSD DiskANN/AISAQ pipeline, the validator keeps
leadership_claim_allowed=false.
| Metric | KnowWhere C++ | Hanns | vs Native |
|---|---|---|---|
| Graph build (Optimize) | 854.2s | 336.9s | 2.53× faster |
| Index load | 1158.9s | 673.7s | 1.72× faster |
| QPS (c=20, ef=128, k=100) | ~500 | 1,051 | ~2.1× faster |
| QPS (c=80, ef=128, k=100) | ~800 | 1,042 | ~1.3× faster |
| Recall@100 | 0.960 | 0.957 | parity |
| ef | Hanns QPS | Lance QPS | Hanns/Lance |
|---|---|---|---|
| 50 | 2,331 | 1,473 | 1.58× |
| 200 | 794 | 483 | 1.64× |
| 800 | 245 | 140 | 1.75× |
Hanns search leads by 1.58–1.75× across the full ef range; advantage grows with higher recall requirements.
| System | Recall@10 | QPS |
|---|---|---|
| DiskANN Rust (Microsoft) | 0.986 | 4,832 |
| Hanns AISAQ NoPQ | 0.994 | 5,806 |
+20% higher QPS, +0.8% better recall at identical parameters on identical hardware.
Standard Product Quantization (PQ) — used by FAISS and most ANN libraries — minimizes L2 reconstruction error. On modern embedding search (high-dim, Inner Product metric), this is the wrong objective: PQ recall collapses near zero.
USQ (Unit Sphere Quantizer) applies a QR orthogonal rotation before quantizing, making compression metric-agnostic. Derived from RabitQ principles and extended to multi-bit (1/4/8-bit) with:
- QR decomposition rotation matrix (trained once, applied to all vectors)
- Unified 1/4/8-bit scalar quantization in the rotated space
- AVX512VNNI integer dot product scoring (
vpdpbusd) - Two-stage pipeline: 1-bit FastScan filter → B-bit rerank
Dataset: Cohere Wikipedia-1M (768-dim, Inner Product), nprobe=32, x86 authority.
| Method | Compression | QPS | Recall@10 | Usable? |
|---|---|---|---|---|
| IVF-PQ m=32 | 8× | 723 | 0.066 | ✗ |
| IVF-PQ m=48 | 5.3× | 502 | 0.127 | ✗ |
| IVF-Flat | 1× | 339 | 0.798 | ✓ |
| IVF-SQ8 | 4× | 605 | 0.805 | ✓ |
| USQ 4-bit | 8× | 1,308 | 0.879 | ✓ |
| USQ 8-bit | 4× | 1,011 | 0.968 | ✓ |
At the same 8× compression ratio where PQ gives recall=0.066, USQ gives 0.879 — a 13× improvement. USQ 8-bit simultaneously delivers 3× faster QPS than IVF-Flat with +17% better recall at ¼ the memory.
On 3072-dim embeddings (SimpleWiki-OpenAI-260K), USQ 8× still achieves recall 0.925 at 1,607 QPS.
Hanns ships as three distinct integration surfaces from a single codebase:
crate-type = ["staticlib", "cdylib", "rlib"]
│ │ │
Milvus C ABI Python/JNI HannsDB / Lance (native Rust)
Integration method: src/ffi.rs exposes 31 #[no_mangle] extern "C" functions that mirror the KnowWhere C++ API exactly (knowhere_create_index, knowhere_add_index, knowhere_search, knowhere_search_with_bitset, …). Milvus links against the compiled staticlib/cdylib with zero source changes — same data format, same query semantics, same bitset filtering interface.
Milvus standalone
└─ KnowWhere shim (C++)
└─ dlopen libhanns.so ← same ABI as native KnowWhere
└─ IndexWrapper::dispatch → HnswIndex / IvfSq8Index / PQFlashIndex / …
The IndexWrapper is the central dispatch struct — a single IndexKind enum that routes all C API calls to the correct Rust index type at zero overhead (no vtable, no allocation on the hot path).
VectorDBBench end-to-end results confirm 2× QPS across HNSW, IVF-Flat, IVF-SQ8, and IVF-PQ under realistic concurrent load:
| Round | Change | QPS (c=80) |
|---|---|---|
| R4 | FFI lazy bitset allocation | 349 |
| R7 | Private rayon ThreadPool (HNSW_NQ_POOL) | 540 |
| R8 | Eliminate BinaryHeap clone + pre-alloc output buffer | 1,042 |
Integration method: HannsDB adds Hanns as a Cargo dependency and uses index types directly:
use hanns::{HnswIndex, IvfFlatIndex, IvfUsqIndex};No FFI boundary, no serialization overhead. HannsDB wraps index instances in its own VectorIndexBackend trait and handles persistence, collection management, and the VectorDBBench client layer. VectorDBBench authority results (x86, 1536-dim, k=100):
| Metric | Result |
|---|---|
| Load | 148.0s |
| Optimize | 87.9s |
| p99 latency | 1.8ms |
| Recall | 0.9756 |
Previously, cosine search p99 reached 110ms due to per-query allocations. After fixing TLS scratch buffer reuse: p99 = 3.5ms (31× improvement).
Integration method: Lance defines an IvfSubIndex trait for pluggable ANN backends. Hanns implements it for HnswIndex, letting Lance's vector lake query engine call Hanns search directly in-process:
// in Lance repo
impl IvfSubIndex for hanns::HnswIndex { … }No serialization boundary, no IPC. Geometry-mean speedup across ef=50–800: 1.64× with equivalent recall. Build is ~26% slower (Lance uses rayon parallel build; Hanns parallel build improvement is in progress).
Lance also wires Hanns-backed IVF_USQ behind cfg(feature = "hanns") in the vector builder / IVF stack, which is what powers the 1B large-top (k=10,000) runs below.
1B large-top capability (PCA-512, cosine, 5 × 194M shards, x86 ECS):
| Tier | Hanns-backed Lance config | Recall@10K | Mean latency | What it demonstrates |
|---|---|---|---|---|
| DRAM | IVF_USQ 4-bit np=256 rf=2 |
0.934 | 957ms | Best sub-0.95 local-storage point; slightly faster than Lance IVF_RQ at the same recall band |
| DRAM | IVF_USQ 8-bit np=256 rf=1 |
0.973 | 1,826ms | Practical high-recall point, but the 8-bit auxiliary index is too large to be a clear DRAM default |
| SSD | IVF_USQ 4-bit np=256 rf=2 |
0.934 | 964ms | Near-DRAM latency on cold NVMe |
| OBS | IVF_USQ 4-bit np=256 rf=1 |
0.803 | 6,555ms | Lowest-latency remote-storage frontier point |
| OBS | IVF_USQ 8-bit np=256 rf=1 |
0.973 | 7,710ms | Collapses the old 0.81-0.96 OBS gap; faster than Lance IVF_SQ (10,518ms) and much faster than Lance IVF_RQ rf=2 (19,306ms) while also improving recall |
| OBS | IVF_USQ 8-bit np=1024 rf=1 |
0.9949 | 21,312ms | Best sub-0.995 remote point |
At the extreme high-recall end, Lance + Hanns reaches 0.9967 recall@10K on OBS at np=1024 rf=2 in 27,988ms. Build is already production-scale as well: the 5-shard IVF_USQ 8-bit index finishes in about 57 min with 117GB auxiliary.idx per shard (about 1.3TB peak temp space).
The practical reading of these numbers is simple:
- On OBS / remote object storage, Hanns-backed
IVF_USQis the key Lance large-top enabler. - On DRAM / SSD, 4-bit USQ gives Lance strong sub-0.95 points, while 8-bit USQ trades much larger index size for higher recall.
Python (feature = "python"): PyO3 bindings in src/python/ expose HnswIndex, IvfFlatIndex, IvfPqIndex, IvfUsqIndex, MemIndex as native Python classes.
JVM / Android (feature = "jni-bindings"): JNI bindings in src/jni/ expose the same index set via @NativeMethod for Java/Kotlin callers.
All benchmarks run on dedicated x86 server,
target-cpu=native. Apple Silicon builds are for fast iteration only — not used as authority evidence.
| Index | Status | Recall@10 | Notes |
|---|---|---|---|
| HNSW | ✅ Leading | 0.972 (SIFT-1M) | +11.9% vs FAISS 8T; 2.53× faster build in Milvus |
| HNSW-SQ | ✅ Ready | 0.992 | Integer precomputed ADC path |
| IVF-Flat | ✅ Leading | 0.978 (SIFT-1M) | 5.2× faster than FAISS 8T (batch parallel) |
| IVF-SQ8 | ✅ Leading | 0.958 (SIFT-1M) | 1.42× faster than FAISS 8T; AVX2 fused decode |
| IVF-USQ | ✅ Ready | 0.905–0.968 (Cohere-1M) | AVX512VNNI; unified 1/4/8-bit |
| IVF-PQ | ✅ Ready | varies | m=32: 0.720 on synthetic data |
| AISAQ (DiskANN Flash) | ✅ Ready | 0.994 NoPQ (SIFT-1M) | On-demand pread + io_uring; 6.5× faster build than native |
| ScaNN | ✅ Ready | 0.969 | Exceeds 0.95 gate at reorder_k=1600 |
| Sparse / WAND | ✅ Ready | 1.0 | Sparse vector retrieval |
| Binary | ✅ Ready | — | Hamming distance |
src/quantization/
usq/ UsqQuantizer — QR rotation + unified 1/4/8-bit quantization
rotator.rs QR decomposition rotation matrix
quantizer.rs training + SIMD scoring (AVX512VNNI)
fastscan.rs AVX512 fast scan (1-bit stage) + topk
searcher.rs two-stage coarse filter + rerank
pq/ Product Quantizer — parallel k-means
sq/ Scalar Quantizer — SQ8/SQ4
pca/ PCA transform — nalgebra SVD (pure Rust, no BLAS)
cargo build --release # LTO + codegen-units=1 + target-cpu=native
cargo test
cargo run --example benchmark --release.cargo/config.toml enables target-cpu=native on x86_64 and aarch64 automatically.
src/
faiss/ core index implementations
quantization/ quantization subsystem (USQ, PQ, SQ, PCA)
ffi/ FFI layer (KnowWhere-compatible ABI)
tests/ integration and regression tests
examples/ full benchmark harness
assets/benchmarks/ comparison charts
docs/ design docs, performance audits
benchmark_results/ authority verdict artifacts (JSON)
scripts/ chart generation, remote build/test
wiki/ operational runbooks and authority numbers
| Dataset | Dim | Metric | Size |
|---|---|---|---|
| SIFT-1M | 128 | L2 | 1M vectors |
| Cohere Wikipedia-1M | 768 | IP | 1M vectors |
| Cohere-1M | 1024 | L2 | 1M vectors |
| SimpleWiki-OpenAI-260K | 3072 | IP | 260K vectors |
Performance numbers are produced on a dedicated x86 server with target-cpu=native (specs). Apple Silicon builds are for fast iteration and pre-screening only.



