WHAT I LEARNED FROM MY STUPID STUFF
I built something. I broke something. I learned something. Here is the evidence. Read this so you don't have to suffer the way I did.
WHAT I LEARNED FROM MY STUPID STUFF
I built something. I broke something. I learned something.
Here is the evidence.
Read this so you don't have to suffer the way I did.
Sockets, concurrency, memory allocators, compilers, and OS boundaries.
Attention, tokenization, quantization, KV caching, and edge inference.
Analyzing performance, latency, microservices vs monoliths, and building to understand.
A Larger Context Window Does Not Remove the Need for RAG
A large context window sounds like it should make retrieval unnecessary. But after working with constrained local models and explicit retrieval pipelines, I started looking at the problem differently: context capacity and context quality are two different engineering problems.
Why Does Segmentation Need Multiple Scales?
I started looking at segmentation architectures from a different angle: the real problem is not simply predicting a class for every pixel, but preserving enough spatial information while building representations with a sufficiently large receptive field.
What Actually Happens Inside an Image Segmentation Model?
An image segmentation model does not simply 'recognize objects'. Mathematically, it transforms a tensor of pixels through a sequence of learned functions and produces a probability distribution for every pixel. This article derives the mathematics behind that process.
BLOG: What Actually Happens When You Call a Framework API?
We use frameworks every day, but how much of their internal machinery do we actually understand? Let's go underneath the API and trace what really happens when a framework processes a request.
What If AI Never Learned Anything? Can Pure Mathematics Build an Intelligent System?
Before neural networks dominated AI, intelligent behavior was already being built from logic, probability, search, optimization, graphs, and mathematical models. But if we completely remove training and build an AI system using only mathematics, what can it actually solve—and where does it fundamentally break?
Why Can Adding an Index Make Your Database Slower?
Indexes make database reads faster—or at least that's what we usually learn. But in production systems, adding an index can actually make writes slower, increase storage pressure, and sometimes even cause the optimizer to choose a worse execution plan.
BLOG: Why Does Adding More CPU Sometimes Make a Backend Slower?
A senior-level deep dive into concurrency, contention, queueing, lock amplification, and why throwing more CPU at a backend can sometimes make the system slower.
BLOG: Why Adding More Threads Can Make Your Backend Slower
A deep dive into concurrency, contention, context switching, queueing, and why increasing the number of workers doesn't necessarily increase throughput.
BLOG: Why Does RAG Sometimes Make an LLM Worse?
RAG is often presented as the obvious solution to hallucination. But retrieval can also make an LLM less accurate, less confident, and sometimes completely wrong. Let's understand why.
Beyond the Abstraction: Notes on Technology, Engineering, and Everything In Between
A personal manifesto and space for exploring technology, systems, AI mechanics, and the ideas that sit beneath the abstractions we use every day.
What Actually Happens Inside a Database Query? From SQL to the Execution Plan
A deeper investigation into how a relational database transforms SQL into an executable plan, covering parsing, binding, optimization, cardinality estimation, index selection, physical operators, and execution.
What Actually Happens Inside an ORM? From Model Definition to SQL Execution
A technical investigation into what really happens between an ORM query and the database, exploring model metadata, query construction, parameter binding, SQL generation, execution, and result hydration.
What Actually Happens Inside a React Render? From State Update to DOM Commit
A technical investigation into the internal lifecycle of a React update, from state mutation and Fiber scheduling to reconciliation and the final commit into the DOM.
What Actually Happens When You Type a URL Into a Browser?
A systems-level investigation into the chain of events triggered by a URL: parsing, DNS resolution, connection establishment, TLS negotiation, HTTP, and browser rendering.
Why Databases Use B-Trees: The Mathematics Behind Efficient Disk-Based Search
A technical investigation into why database indexes rely on B-Trees and B+Trees, and how high fan-out transforms disk-based search from an expensive linear process into logarithmic traversal.
Distributed Data Integrity: Patterns for Microservices Architecture
A technical deep dive into maintaining data integrity across microservices, exploring Distributed Transactions, the Saga pattern, Transactional Outbox, and the shift toward Eventual Consistency.
Next.js Evolution: The Paradigm Shift from Pages Router to App Router
An architectural analysis of the monumental shift from the legacy Pages Router to the modern App Router in Next.js, highlighting RSCs, routing conventions, and data fetching paradigms.
Demystifying Next.js Hydration and the 'Hydration Mismatch' Error
An in-depth look into the mechanics of Next.js Hydration, why the infamous 'Hydration Mismatch' error occurs, and the standard engineering practices to resolve it.
Next.js App Router: The Architectural Divide Between Server and Client Components
An architectural breakdown of React Server Components (RSC) in the Next.js App Router, exploring the strict divide between Server and Client execution environments.
React Lifecycle Mechanics: Mapping Component Synchronization to useEffect
A deep dive into the React component lifecycle, exploring the mental shift from class-based lifecycle methods to functional synchronization using the useEffect hook.
State Management in React: Local, Global, and the Server State Paradigm
A comprehensive guide to drawing boundaries between Local, Global, and Server State in modern React, featuring tools like Zustand, Redux Toolkit, and TanStack Query.
Demystifying Rendering: Server-Side (SSR) vs. Client-Side (CSR) in Next.js
A comprehensive breakdown of Server-Side Rendering (SSR) and Client-Side Rendering (CSR), their architectural trade-offs, and how to choose the right pattern for your Next.js applications.
Autoregressive Masking: Formalizing the Causal Mask in Transformer Decoder Architectures
A technical analysis of the Causal Mask, the structural constraint that enforces autoregressive generation in Decoder-only Transformers. This paper derives the mechanism's mathematical basis and its role in preventing attention leakage across future tokens.
Inverse Diffusion and Latent Manifolds: Formalizing Generative Mechanics in AI Synthesis
An investigation into the mathematical foundations of Generative AI, focusing on Inverse Diffusion processes and Latent Space formalization for image and text synthesis.
Multi-Head Attention: The Engine of Parallel Representation in Transformers
A comprehensive breakdown of Multi-Head Attention, the mathematical framework that allows Transformers to capture parallel semantic subspaces simultaneously.
Understanding the Context Window: The Short-Term Memory of LLMs
An exploration of the context window in Large Language Models, detailing its token-based architecture, O(N^2) computational complexity, and the 'Lost in the Middle' phenomenon.
Scaling in Transformer Architectures: The Mathematical Rationale behind $\sqrt{d_k}$
A derivation of the variance explosion in high-dimensional dot products and its deleterious effects on softmax saturation and gradient propagation.
Statistical Tokenization: Formalizing the Byte Pair Encoding (BPE) Algorithm for Subword Decomposition
A technical investigation into Byte Pair Encoding (BPE), the subword tokenization standard for Large Language Models. This paper details the iterative transition from character-level granularity to high-density subword dictionaries.
Attention Dynamics: Formalizing Scaled Dot-Product Mechanisms in Transformer Architectures
A technical formalization of the Scaled Dot-Product Attention mechanism. This paper analyzes the topological interaction between Queries, Keys, and Values, providing a step-by-step numerical derivation of the attention pipeline.
Synthesizing Minority Samples: A Formal Analysis of Linear Interpolation in Imbalanced Classification
A rigorous mathematical investigation into the Synthetic Minority Over-sampling Technique (SMOTE). This paper details the k-NN selection process and the geometric foundations of linear interpolation used to expand decision boundaries in imbalanced datasets.
Strategic Informatics: A Formal Investigation into Undersampling Mechanisms for Imbalanced Classification
A technical exploration of majority class reduction strategies. This paper formalizes Random Undersampling, the NearMiss heuristic suite, and Tomek Link boundary cleaning for optimizing inference in high-imbalance network traffic datasets.
Automata as Memory: Decoding LSTM State Persistence in Terminal Sequences
A rigorous mathematical analysis of the divergence between Cell State ($C_L$) and Hidden State ($h_L$) at the terminal step of Long Short-Term Memory architectures. Explores the functional roles of these states in Many-to-One and Many-to-Many topologies.
Embedding Vector vs Standard Vector: The Mathematical Soul of Modern AI
A comparative study between engineered standard vectors and learned embedding vectors, exploring latent feature spaces and semantic arithmetic in Deep Learning.
The Calculus of Compression: Mathematical Foundations of Post-Training Quantization (PTQ)
A formal exploration of affine quantization mapping. This paper details the derivation of scaling factors and zero-points for converting FP32 tensors to INT8 precision while preserving structural fidelity during inference.
Foundations of Recurrent Architectures: Parameter Sharing and Temporal Dynamics
An analytical study of Recurrent Neural Networks (RNNs), examining the mathematical mechanics of parameter sharing, temporal hidden states, and the vanishing gradient bottleneck.
LoRA vs QLoRA: The Ultimate Memory Bottleneck Showdown
A deep dive comparing LoRA and QLoRA, analyzing their mathematical mechanics, memory constraints, and how they democratize LLM fine-tuning.
Mathematical Foundations of Spatial and Temporal Subsampling: A Study on Pooling Layers
A rigorous mathematical analysis of dimensionality reduction in Deep Learning, exploring the formal mechanics of Max, Average, and Global Pooling across 1D, 2D, and 3D architectures.
Taxonomy of Machine Learning Optimization: A Survey of Training Paradigms
A systematic categorization of algorithmic training methodologies in Artificial Intelligence, analyzing the mathematical foundations of Supervised, Unsupervised, and Reinforcement Learning.
Mathematical Foundations of Convolutional Architectures: A Spatiotemporal Research Study
A rigorous mathematical exploration of convolutional operations in 1D, 2D, and 3D spaces. Analyzing receptive field dynamics, computational complexity, and dimensionality mapping in deep neural networks.
Distributed Data Parallel (DDP) Architecture: Mathematical Foundations of Ring All-Reduce
A rigorous mathematical and architectural analysis of Distributed Data Parallel (DDP) in PyTorch. Explores the GIL bottlenecks of legacy systems and the efficiency of the Multi-process Ring All-Reduce topology.
Computational Efficiency in Edge AI: Optimization via Pythonic Lazy Evaluation
An analysis of memory management strategies for resource-constrained Edge AI devices, focusing on the mechanics of Python Generators and the 'yield' primitive.