Abstract
Approximate Nearest Neighbor (ANN) search is a foundational primitive in modern recommendation and retrieval systems. However, production deployments routinely require filtered ANN search-retrieving the top- K nearest neighbors that also satisfy metadata predicates-a requirement that popular highperformance libraries such as FAISS and ScaNN do not natively support without costly post-filtering loops. Existing solutions that support filtering, such as distributed search engines, introduce unacceptable latency and infrastructure overhead for latencysensitive candidate retrieval pipelines. We present ANNex, a production ANN system that integrates metadata filtering directly into HNSW graph traversal, eliminating the need for post-filtering iteration. ANNex introduces a Decreasing- K traversal strategy-the inverse of post-filtering's Increasing- K loop-in which already-visited nodes are tracked and excluded from subsequent traversals, reducing graph search depth with each iteration rather than increasing it. Combined with Product Quantization for memory compression, integer key optimization, and compiled filter functions for sub-millisecond predicate evaluation, ANNex achieves sub-30 ms p99 latency at production scale on 10 million 512-dimensional vectors with up to four concurrent clients per instance. We describe the system architecture, key design tradeoffs, and empirical evaluation results, providing a practical reference for practitioners building filtered ANN systems at scale.


