pomaidb
pomagrenate/pomaidb
Overview
A predictable, embedded multimodal vector database and offline RAG engine for Edge AI and resource-constrained devices. Built in C++20 with a zero-OOM memory model, append-only flash storage, and an integrated arena memory allocator.
Technologies & Concepts
Project Analysis
🔴 Problem
Cloud-first assumptions on edge devices: Most existing vector databases are built for large cloud servers with tens of gigabytes of RAM, multi-core CPU topologies, and high-end NVMe storage. When deployed on embedded devices such as Raspberry Pi, Orange Pi, or IoT gateways (which often have 512 MB to 2 GB RAM), they frequently crash due to uncontrolled memory growth and Out-Of-Memory (OOM) killer invocations.
Destructive flash drive wear: Consumer edge devices rely on MicroSD cards or raw eMMC flash memory, which have limited write endurance. Storage engines that rely on frequent in-place updates or heavy background B-tree/LSM compactions suffer from severe Write Amplification Factors (often 5x to 10x), rapidly degrading flash hardware and causing disk I/O stalls.
Tail latency jitter & thread contention: In real-time edge workloads such as camera gateways or local robotics, unpredictable garbage collection cycles and multithreaded mutex contention introduce large tail latency spikes (p99 and p99.9), making deterministic real-time response times difficult to guarantee.
Heap fragmentation under continuous influx: High-frequency vector ingestion and deletion cycles cause standard OS heap allocators (glibc malloc or MSVCRT) to suffer from severe external fragmentation, gradually inflating the Resident Set Size (RSS) even when logical data sizes remain constant.
⚪ Baseline
Starting point: Standard embedded vector libraries and databases (such as SQLite with vector search extensions, embedded ChromaDB instances, or raw FAISS/HNSW index graphs) running directly inside resource-constrained environments.
Baseline characteristics:
- Unbounded memory consumption during batch streaming ingestion, leading to process termination by the operating system
- Frequent random in-place writes and high write amplification factors (5.8x to 9.4x), resulting in rapid MicroSD endurance degradation
- Multi-threaded synchronization overhead on low-power, low-core ARM architectures
- Standard system memory allocation with 25% to 30% heap fragmentation under adversarial vector churn
- Complex external infrastructure requirements and lack of native zero-copy multimodal abstractions
Expected trade-off: Inability to sustain continuous streaming vector workloads within strict 32 MB to 64 MB hardware memory envelopes without risking OOM termination or storage wear out.
🔵 Change
Built PomaiDB in C++20: Designed an embedded, multimodal vector database and offline RAG engine centered around a deliberate principle: one process, one logical database, and one predictable execution path.
Key architectural changes:
- Pomegranate storage anatomy: Implemented a two-tier storage hierarchy: an in-memory active Rind (MemTable) for immediate vector insert/taste, flushed by an asynchronous Press worker into immutable on-disk Locules (segments) containing Arils (clusters), Seed Kernels (quantized/full vectors), Seed Scars (metadata filters), and Pulp (navigational graph connectivity).
- Deterministic zero-OOM memory bounding: Integrated an active
auto_freeze_on_pressuremechanism with configurable RSS caps (e.g. 32 MB or 64 MB). When memory pressure is detected, PomaiDB automatically flushes saturated Rinds to disk rather than allowing memory usage to expand indefinitely. - Native palloc memory backbone: Replaced standard heap allocators on the critical vector path with
palloc, an arena slab allocator providing guaranteed 64-byte alignment for AVX2/AVX-512 and ARM NEON SIMD operations, with O(1) batch reset capabilities. - Multi-format quantization with SeedKernel reranking: Implemented Float32, FP16, Int8 (SQ8 Pulp), and 1-Bit (Binary Quantization) modes. Fast bitwise Hamming distance and integer dot products scan candidate vectors, followed by exact SeedKernel reranking to preserve accuracy.
- Flash-friendly log-structured I/O: Enforced append-only sequential writes and tombstone-based deletions for immutable Locules, achieving a near-ideal 1.05x Write Amplification Factor on flash and MicroSD media.
- Single-threaded event loop: Eliminated lock contention, mutex-heavy synchronization, and race conditions across the hot query and ingestion paths.
🟣 Measurement
Test environment: HP ProBook 450 G5, Intel Core i7-8550U (8 logical threads @ 1.80GHz), 16 GB RAM, SATA SSD, alongside simulated memory-constrained container environments (32 MB to 128 MB RAM ceilings).
Benchmark suites:
ci_perf_bench.cc: In-memory Rind insertion and taste queries for 64-dimensional vectors under CI gatescomprehensive_bench.cc: 10,000 vectors at 128 dimensions across 1,000 queries with Top-k=10 searchquantization_comp_bench.cc: Comparative evaluation of FP32, FP16, SQ8, and 1-Bit BQ across 10,000 vectors at 256 dimensionspalloc_env_stress.cc: Microarchitectural hardware PMU metrics, SIMD throughput, cache lines, and TLB pressurebenchmark_a (Stress Suite): Multi-environment validation across IoT Starvation (5k vecs @ 1536-dim), Edge Churn (10k vecs over 5 cycles with leak checking), and Cloud Scale (20k vecs @ 1536-dim with 3 disk flushes)
Metrics collected: Ingestion throughput (vectors/sec), latency per vector (µs), query latency percentiles (p50, p90, p95, p99, p99.9), Recall@10 accuracy, resident set size (RSS in MB), Write Amplification Factor (WAF), and data integrity verification.
🟢 Result
Microsecond ingestion latency: Ingestion latency dropped to 8.23 µs per vector in batch mode (up to 121,528 vectors/sec at 128 dimensions) and 10.48 µs for individual single puts.
Deterministic tail latency bounds: Sub-25 ms p50 latency across evaluated dimensions (2.13 ms for 64-dim in-memory Rind, 18.45 ms for 256-dim Int8 SQ8) with a tight p99.9 of 34.10 ms on 256-dim SQ8, completely avoiding garbage collection jitter.
Pareto-efficient quantization: 100.0% Recall@10 preserved across Float32 down to 1-Bit Binary Quantization through SeedKernel reranking, while reducing in-memory size by up to 32x (0.32 MB vs 10.24 MB for 10,000 vectors).
Deterministic zero-OOM memory bounding: Under a continuous 100,000-vector streaming workload, active RSS remained strictly bounded under the configured 32 MB threshold by automatically freezing Rinds to immutable Locules, whereas unconstrained allocation grew past 120 MB.
Minimal write amplification: Append-only Locule writes achieved a Write Amplification Factor of 1.05x, compared to 7.85x for naive full rewrites, protecting consumer flash and MicroSD cards from premature endurance failure.
Resilience under high-dimensional stress: 100% data integrity verified with zero memory leaks across 5 consecutive restart cycles on massive 1536-dimensional embeddings (6 KiB/vector).
Performance Visualizations & Empirical Charts
1. Pomegranate Engine Architecture & Benchmark Overview

Comprehensive summary showing ingestion throughput across dimensions (up to 131,904 vec/s at 64-dim), 100% Recall@10 across precision modes, microsecond-scale vector ingestion latency (8.23 µs batch), and storage footprint reduction down to 1.6 MB per 100k vectors with 1-Bit BQ.
2. Memory Backbone: Palloc Microarchitectural Performance

Empirical evaluation of the integrated palloc allocator: 8.4x to 60.53x peak speedup in vector batch churn, 6.8x lower p50 allocation latency (14 ns vs 95 ns), 3.57x–4.10x SIMD throughput speedup, and a 33% reduction in TLB pressure over the OS allocator.
3. Ingestion Engine: Mode Scaling & Throughput Profile

Ingestion throughput and per-vector latency across single puts (95,408 vec/s, 10.48 µs), sequential WAL appends (117,705 vec/s, 8.50 µs), and tuned batch sizes (up to 121,528 vec/s at 1k batch size) measured across 100,000 vectors at 128 dimensions.
4. Quantization Pareto Frontier: Memory Compression vs Accuracy

Pareto frontier comparing Float32 (10.24 MB), FP16 (5.12 MB), Int8 SQ8 (2.56 MB), and 1-Bit BQ (0.32 MB) across 10,000 vectors at 256 dimensions. SeedKernel reranking enables aggressive 32x compression while sustaining 100.0% Recall@10.
5. Search Latency Percentiles & Tail Predictability

Logarithmic latency distribution across percentiles (p50 to p99.9). Single-threaded execution eliminates mutex contention, maintaining sub-millisecond p50 on in-memory Rinds and a tight 34.10 ms p99.9 on 256-dim Int8 SQ8 searches.
6. Deterministic Memory Bounding: Auto-Freeze Mechanism

Continuous 100,000-vector ingestion workload under a 32 MB pressure cap. While an unconstrained memtable accumulates memory past 120 MB (risking container OOM kills), PomaiDB's auto-freeze mechanism deterministically resets the active Rind into immutable Locules.
7. Storage Longevity: Append-Only Locule Press vs Full Rewrite

Write Amplification Factor (WAF) comparison showing PomaiDB's O(1) append-only Press achieving 1.05x WAF versus 7.85x for naive segment rewrites, significantly extending flash and MicroSD drive lifespan on edge hardware.
8. Multi-Environment Stress Benchmark (benchmark_a Suite)

Validation under three edge conditions using 1536-dimensional embeddings: IoT Starvation (pure in-memory, 1,691 vec/s), Edge Churn (5 restart cycles with zero memory leaks, 1,455 vec/s), and Cloud Scale (sustained with 3 disk flushes, 1,078 vec/s). 100% data integrity verified.
🟡 Lesson
Edge constraints forced better architectural choices: When I designed for small devices with 64 MB to 128 MB of RAM, I couldn't rely on bloated distributed abstractions or lazy memory reclamation. Constraining the operating environment forced me to think carefully about every memory allocation, pointer dereference, and disk write.
The single-threaded execution model paid off: I initially wondered whether avoiding multi-threading would bottleneck query throughput. What I noticed in practice was that eliminating mutex locks and thread synchronizations kept CPU caches hot and eliminated lock contention entirely. The result was predictable microsecond-level ingestion without the random latency spikes typical of multithreaded systems.
Flash storage endurance is an overlooked constraint: When I first started building PomaiDB, my focus was primarily on vector search speed. However, after investigating how embedded storage works, I realized that high write amplification destroys consumer flash memory within months. Switching to an append-only, log-structured model with tombstone deletions brought the Write Amplification Factor down to 1.05x, which is critical for edge device longevity.
Quantization with SeedKernel reranking exceeded my expectations: I initially assumed that 1-bit binary quantization would degrade retrieval accuracy too severely for practical RAG workloads. What surprised me was that combining fast bitwise candidate filtering with full-precision SeedKernel reranking achieved 100.0% Recall@10 while reducing memory consumption by 32x.
Memory bounding requires active backpressure: I learned that passive memory tracking is insufficient on edge hardware. Without active auto-freezing and ingestion backpressure, sudden bursts of incoming data will quickly breach physical memory limits and trigger the OS OOM killer.
What I still want to investigate: I want to test PomaiDB across a wider variety of physical ARM64 boards (such as the Raspberry Pi Zero 2W and Orange Pi Zero 3), explore writing dedicated ARM NEON vector assembly kernels for distance calculation, and test how the embedded RAG pipeline behaves when paired directly with small quantized local LLMs in offline environments.