palloc
pomagrenate/palloc
Overview
An ultra-fast, lightweight, and thread-safe general-purpose memory allocator. Drop-in malloc replacement built for high throughput and low latency.
Technologies & Concepts
Project Analysis
🔴 Problem
mimalloc limitation: Microsoft's mimalloc is an excellent general-purpose allocator, but it's not optimized for vector/embedding workloads where memory allocations are contiguous and huge.
Vector allocation patterns: Vector operations in machine learning and embedding workloads require large, contiguous memory blocks that are allocated and freed in batch patterns, which general-purpose allocators don't handle optimally.
Performance bottleneck: The overhead of individual malloc/free operations for large contiguous blocks creates significant performance penalties in high-throughput vector processing scenarios.
⚪ Baseline
Starting point: mimalloc (Microsoft Research) - a high-performance general-purpose memory allocator
Baseline characteristics: Excellent for general workloads, but not optimized for:
- Large contiguous memory blocks
- Batch allocation/deallocation patterns
- Vector and embedding workloads
- SIMD-aligned memory access
Expected performance: Standard malloc/free overhead for large allocations
🔵 Change
Forked mimalloc: Created palloc as a specialized fork of mimalloc optimized for vector/embedding workloads
Key architectural changes:
- Arena-based allocation with O(1) reset vs O(N) individual frees
- 64-byte alignment guarantees for SIMD operations (AVX-512)
- Contiguous memory allocation for better TLB utilization
- Reduced fragmentation in batch allocation scenarios
- Optimized for large, contiguous memory blocks
Implementation approach: Drop-in malloc replacement via LD_PRELOAD/DYLD_INSERT_LIBRARIES/Windows DLL redirect
🟣 Measurement
Test environment: Linux with POSIX override, Dell Latitude E5440, Intel Core i5 (2 cores), 8GB RAM
Benchmark suite: Adversarial benchmarks for vector/embedding workloads including:
- Vector batch churn (batch sizes: 32, 128, 512, 4096)
- SIMD latency tests (hot/cold cache, aligned/unaligned)
- TLB pressure tests (working sets up to 8GB)
- Comparison: palloc vs system allocator
Metrics collected: Throughput, latency, memory usage, cache performance, TLB efficiency
🟢 Result
Exceptional batch performance: 18.4x average speedup, 60.53x peak speedup in batch churn scenarios
SIMD improvements: 3.57x speedup (hot aligned), 4.10x speedup (cold unaligned) for memory operations
Latency reductions: p50 latency reduced from 37-165ns (system) to 6-37ns (palloc)
TLB benefits: 1.33x improvement in TLB pressure tests with large working sets
Exceeded expectations: Significantly outperformed expected 2-4x batch improvements, achieving 18.4x average
Performance Visualizations
Speedup Comparison

Throughput Comparison

Latency Distribution

Memory Usage Analysis

🟡 Lesson
Architecture matters for specific workloads: General-purpose allocators like mimalloc are excellent, but specialized allocators can provide order-of-magnitude improvements for specific workload patterns.
Arena-based allocation power: The O(1) reset operation vs O(N) individual frees creates dramatic performance differences at scale, especially for batch allocation patterns common in ML workloads.
Memory alignment benefits: Guaranteed 64-byte alignment provided measurable SIMD benefits (3.57x-4.10x), validating the design decision to prioritize alignment for vector operations.
Benchmark design complexity: I learned that benchmarking memory allocators is non-trivial. Workload patterns, warm-up iterations, and measurement methodology all significantly impact results. The adversarial benchmark suite was crucial for testing real-world vector/embedding workloads.
Performance vs expectations: I initially expected 2-4x improvements for batch operations, but achieved 18.4x average with 60.53x peak. This taught me that well-designed specialized allocators can outperform general-purpose ones by orders of magnitude for specific workloads.