Skip to main content
←Projects
Side Projects
PK

pomaikache

pomagrenate/pomaikache

Overview

An ultra-high-throughput, in-memory vector cache and ingestion engine built exclusively for dense floating-point embeddings in pure C. Designed with contiguous aligned memory layouts, zero-copy lock-free ring buffers, and SIMD acceleration to achieve 5.9M+ ops/sec throughput, sub-3ms tail latency, and >99% recall without generic key-value overhead.

Benchmark Host Device & Environment

All benchmarks and latency profiles were measured live on the following hardware platform without synthetic virtualization:

Processor
Intel(R) Core(TM) i5-4310U CPU @ 2.00GHz (Haswell, Max 2.60 GHz)
Cores & Threads
2 Physical Cores, 4 Logical Processors
Memory (RAM)
8.00 GB DDR3
Operating System
Microsoft Windows 10 Pro (64-bit, Build 19045)
Vector Extension
AVX2, FMA3, SSE4.2

Key Benchmark Measurements

Peak Ingestion (Ring Buffer)
5,919,415 ops/s
Lock-free ring buffer pipeline
Multi-Thread Throughput (16-Thr)
1,199,798 ops/s
Under sustained concurrency stress
Search Median Latency (p50)
1,537.07 µs
~1.54 ms median response time
Search Tail Latency (p95)
2,928.98 µs
~2.93 ms 95th percentile tail
Search High Tail (p99)
12,607.11 µs
12.61 ms 99th percentile spike
Search Recall@10 (64d - 128d)
100.0%
Zero accuracy loss on low dimensions
Search Recall@10 (768d)
99.7%
Tested on 768d embedding models
Search Recall@10 (1536d)
99.6%
High-dim OpenAI ada/v3 embeddings

Technologies & Concepts

CVector CacheDense EmbeddingsLock-Free Ring BufferSIMD / AVX2High Throughput (5.9M ops/s)Microsecond Tail LatencyIn-Memory Database

Project Analysis

🔴 Problem

The mismatch of general-purpose caches for vector embeddings: In high-throughput AI pipelines (such as real-time RAG, semantic caching, and recommendation scoring), systems need to ingest and query millions of dense floating-point vectors every second. Traditional in-memory key-value engines (e.g., Redis, Memcached, Dragonfly) were architected around strings, hashes, and generic keys.

Memory indirection & serialization bottlenecks: When vectors are packed into generic key-value stores, each vector payload suffers from pointer chasing, heap fragmentation, and serialization overhead. Standard Redis with RediSearch vector extensions typically bottlenecks at approximately 85,000 operations per second because the single-threaded event loop and dynamic hash tables cannot take advantage of uniform contiguous vector layouts.

Tail latency spikes under concurrent load: Generic memory caches exhibit severe tail latency spikes when background memory compaction or rehashing collides with incoming high-dimensional vector similarity calculations.

⚪ Baseline

Conventional in-memory systems evaluated:

  • Redis 7.2 (RediSearch vector index): Ingestion throughput capped at ~85,000 ops/s due to single-threaded command processing and generic allocator overhead.
  • Dragonfly 1.14 (Multi-threaded Key-Value): Reached ~450,000 ops/s on standard key-value operations, but lacks specialized contiguous memory vector indexing.
  • Memory layout: Unaligned, dynamic heap-allocated structures requiring pointer dereferencing on every dimension comparison.

🔵 Change

Engineered Pomaikache as a dedicated C vector cache: I designed Pomaikache from scratch in pure C specifically and exclusively for vector operations, stripping away all non-vector overhead.

Key architectural decisions:

  • Lock-Free Ring Buffer Ingestion: Implemented a cache-aligned circular ring buffer with atomic sequence counters, allowing concurrent worker threads to ingest vectors with zero lock contention.
  • Contiguous Aligned Vector Arenas: Enforced 32-byte and 64-byte memory alignment (`posix_memalign` / `_aligned_malloc`) so that vector floats sit contiguously in memory, perfectly matching CPU L1/L2 cache lines and SIMD registers.
  • SIMD AVX2 Inner Product & Distance Kernels: Vector dot products and Euclidean distance computations utilize AVX2 vector intrinsics to process 8 single-precision floats per instruction cycle.
  • Zero-Copy Memory Semantics: Vector ingestion writes directly into pre-allocated memory pools, avoiding intermediate string conversions and JSON/RESP protocol parsing.

🟣 Measurement

Evaluation context: Measured directly on the host machine (Intel Core i5-4310U @ 2.00GHz, 2 Cores/4 Threads, 8GB DDR3 RAM, Windows 10 Pro) using live ingestion generators and search query workloads.

Key metrics evaluated:

  • Throughput Comparison: Measured against Redis 7.2 (RediSearch) and Dragonfly 1.14 under identical vector workloads.
  • Tail Latency Breakdown: Profiled search query latency percentiles (p50, p95, p99) in microseconds across thousands of live requests.
  • Recall Accuracy Across Dimensions: Tested Search Recall@10 against brute-force exact nearest neighbor across 64d, 128d, 384d, 768d, and 1536d vector spaces.

🟢 Result

5.9M+ ops/s peak throughput: Pomaikache achieved 5,919,415 operations per second via its lock-free ring buffer pipeline—a ~70x speedup over Redis 7.2 (85k ops/s) and ~13x over Dragonfly (450k ops/s). Under sustained 16-thread stress loads, it delivered 1,199,798 ops/s.

Microsecond tail latency profile: Recorded a median latency (p50) of 1,537.07 µs (~1.54 ms) and a 95th percentile tail (p95) of 2,928.98 µs (~2.93 ms), demonstrating minimal latency jitter.

Preserved accuracy (>99.2% Recall@10): Delivered 100% recall on 64d and 128d vectors, 99.8% on 384d, 99.7% on 768d, and 99.6% on 1536d high-dimensional embeddings.

Benchmark Visualizations & Performance Evidence

Live Ingestion & Processing Throughput Comparison (Ops / Sec)
Live Ingestion and Processing Throughput Comparison

Throughput comparison contrasting Redis 7.2 RediSearch (85,000 ops/s) and Dragonfly 1.14 (450,000 ops/s) with Pomaikache 1.0 (1,199,798 ops/s under 16-thread stress and 5,919,415 ops/s using the lock-free ring buffer).

Vector Search Tail Latency Profile (µs)
Pomaikache Vector Search Tail Latency Profile

Tail latency distribution highlighting p50 median latency at 1,537.07 µs (~1.54 ms), p95 tail latency at 2,928.98 µs (~2.93 ms), and p99 high tail at 12,607.11 µs (~12.61 ms).

Search Recall@10 Accuracy Across Vector Dimensions (64d to 1536d)
Search Recall@10 Accuracy Across Vector Dimensions

Recall accuracy validation demonstrating >99.2% Recall@10 across varying vector dimensions: 100.0% at 64d/128d, 99.8% at 384d, 99.7% at 768d, and 99.6% at 1536d.

🟡 Lesson

Specialization beats generic abstraction for vector math: Trying to force vector embeddings through general-purpose key-value memory allocators wasted substantial CPU cycles on pointer management. Aligning contiguous memory directly in C showed me that when the memory structure matches the hardware cache line, performance jumps orders of magnitude.

Lock-free ring buffers eliminate thread contention: In multi-threaded ingestion tests, standard mutexes became the primary bottleneck once concurrency scaled beyond 4 threads. Moving to an atomic ring buffer decoupled producers from consumers and unlocked 5.9M ops/s throughput on modest hardware.

Tail latency matters more than mean latency: When designing vector caching for conversational AI and RAG, a low average latency is meaningless if the 99th percentile stalls for hundreds of milliseconds. Ensuring bounded memory allocations kept p95 tail latency under 3 milliseconds.

Explore This Project