<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Prathamesh Blog's]]></title><description><![CDATA[Prathamesh Blog's]]></description><link>https://prathamesh-thakare.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/648c30cf85f011423c0e5c2a/58df8675-9cca-49b3-ab64-730f334cf05e.png</url><title>Prathamesh Blog&apos;s</title><link>https://prathamesh-thakare.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sat, 10 Oct 2026 05:47:07 GMT</lastBuildDate><atom:link href="https://prathamesh-thakare.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Building Rag System From First Principals]]></title><description><![CDATA[What Is RAG?
In general, Large Language Models are incredibly capable. They can answer questions, summarize text, write code, and perform many other tasks.
But they are still general-purpose models. T]]></description><link>https://prathamesh-thakare.hashnode.dev/building-rag-system-from-first-principals</link><guid isPermaLink="true">https://prathamesh-thakare.hashnode.dev/building-rag-system-from-first-principals</guid><dc:creator><![CDATA[Prathamesh Thakare]]></dc:creator><pubDate>Tue, 29 Sep 2026 06:47:50 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/648c30cf85f011423c0e5c2a/1ff83261-748d-4545-88f5-cbe9c30d7226.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What Is RAG?</h2>
<p>In general, Large Language Models are incredibly capable. They can answer questions, summarize text, write code, and perform many other tasks.</p>
<p>But they are still <strong>general-purpose models</strong>. They don't automatically have access to your private or application-specific data.</p>
<p>So, what if you want an LLM to answer questions about <strong>your data</strong>?</p>
<p>For example, imagine you have an internal company document containing:</p>
<blockquote>
<p>"Employees can work remotely for up to 3 days per week."</p>
</blockquote>
<p>Now you ask an LLM:</p>
<blockquote>
<p>"How many days can I work remotely?"</p>
</blockquote>
<p>A general-purpose LLM doesn't automatically have access to that document.</p>
<p>You could provide the document along with the question, but imagine doing this with thousands of documents and hundreds of questions. Sending all that information to the LLM every time would be expensive, inefficient, and eventually run into context limits.</p>
<p>Instead, we can give the LLM access to the information it needs <strong>when it needs it</strong>.</p>
<p>This is where <strong>RAG — Retrieval-Augmented Generation</strong> comes in.</p>
<p>The idea is fairly simple:</p>
<p><strong>User Question → Retrieve relevant information → Give it to the LLM → Generate an answer</strong></p>
<p>For the same question:</p>
<blockquote>
<p>"How many days can I work remotely?"</p>
</blockquote>
<p>The RAG system first searches through the available documents and retrieves the relevant information:</p>
<blockquote>
<p>"Employees can work remotely for up to 3 days per week."</p>
</blockquote>
<p>That information is then passed to the LLM along with the user's question.</p>
<p>The LLM doesn't need to know our company's remote-work policy beforehand. It just needs the relevant context to generate an answer.</p>
<p>And while the basic idea sounds simple, it introduces a few interesting problems.</p>
<p><strong>Let's understand them one by one.</strong></p>
<h2>Key Concepts</h2>
<h3>Chunking</h3>
<p>At this point, our biggest problem is that we have to send the <strong>whole document as context</strong> to the LLM for it to answer our query.</p>
<p>But in most cases, the answer we're looking for might exist in just a small part of that document.</p>
<p>For example, imagine we have a 100-page company policy document, and the user asks:</p>
<blockquote>
<p>"How many days can I work remotely?"</p>
</blockquote>
<p>The answer might be just one sentence somewhere in those 100 pages.</p>
<p>There is no reason to send all 100 pages to the LLM when we only need that small piece of information.</p>
<p>So the first question becomes:</p>
<p><strong>How can we break a large document into smaller, meaningful pieces so that we can later work with only the relevant context?</strong></p>
<p>This process is called <strong>chunking</strong>.</p>
<p>Instead of treating the entire document as one large piece of text, we split it into smaller sections called <strong>chunks</strong>.</p>
<p>For example:</p>
<pre><code class="language-text">Document
│
├── Chunk 1
├── Chunk 2
├── Chunk 3
├── Chunk 4
└── ...
</code></pre>
<p>Now, instead of eventually giving the LLM the entire document, we can aim to give it only the chunks that are relevant to the user's question.</p>
<p>We'll worry about <strong>how to find those relevant chunks</strong> in the retrieval section.</p>
<p>For now, the important idea is:</p>
<blockquote>
<p><strong>We don't need the entire document to answer every question. We need the right part of the document.</strong></p>
</blockquote>
<p>And chunking is the first step towards making that possible.</p>
<h3>Embeddings</h3>
<p>Now we have our document split into smaller chunks.</p>
<p>But we still have a problem.</p>
<p>Suppose we have these chunks:</p>
<p><strong>Chunk 1</strong></p>
<blockquote>
<p>Employees can work remotely for up to 3 days per week.</p>
</blockquote>
<p><strong>Chunk 2</strong></p>
<blockquote>
<p>Employees are eligible for 20 days of annual leave.</p>
</blockquote>
<p><strong>Chunk 3</strong></p>
<blockquote>
<p>Health insurance covers immediate family members.</p>
</blockquote>
<p>And the user asks:</p>
<blockquote>
<p>"How many days can I work from home?"</p>
</blockquote>
<p>How do we know that <strong>Chunk 1</strong> is the relevant one?</p>
<p>A simple keyword search might look for words like <code>work</code>, <code>home</code>, or <code>days</code>. But the question and the chunk don't necessarily have to use the exact same words.</p>
<p>For example:</p>
<blockquote>
<p>"How many days can I work from home?"</p>
</blockquote>
<p>and</p>
<blockquote>
<p>"Employees can work remotely for up to 3 days per week."</p>
</blockquote>
<p>use different words, but they have a similar meaning.</p>
<p>This is where <strong>embeddings</strong> come in.</p>
<p>Essentially, we convert text into a numerical representation called an <strong>embedding</strong>.</p>
<pre><code class="language-text">Chunk 1 → [0.12, -0.45, 0.78, ...]
Chunk 2 → [0.81,  0.21, -0.32, ...]
Chunk 3 → [-0.14, 0.67, 0.19, ...]
</code></pre>
<p>We can do the same thing with the user's question:</p>
<pre><code class="language-text">"How many days can I work from home?"
                ↓
        [0.15, -0.42, 0.74, ...]
</code></pre>
<p>The important idea is that text with similar semantic meaning should have embeddings that are <strong>closer together</strong>.</p>
<p>So the embedding for our question should be closer to <strong>Chunk 1</strong> than to Chunk 2 or Chunk 3.</p>
<p>This allows us to search based on <strong>semantic similarity</strong>, rather than relying only on exact keyword matches.</p>
<p>And this idea of converting text into numerical representations isn't unique to RAG.</p>
<p>At a high level, this is also part of how modern language models work: text is ultimately converted into numerical representations that models can process and operate on.</p>
<p>RAG uses a similar idea for a different purpose — we use embeddings to make our documents <strong>searchable by meaning</strong>.</p>
<p>Now we have:</p>
<p><strong>Document → Chunks → Embeddings</strong></p>
<p>But we still need to compare the user's question against these embeddings and decide which chunks are relevant.</p>
<p>That's where <strong>retrieval</strong> comes in.</p>
<h3>Retrieval</h3>
<p>Now we have our document split into chunks, and each chunk has an embedding.</p>
<p>We also have an embedding for the user's question.</p>
<p>But having these embeddings alone doesn't give us an answer.</p>
<p>We need to actually <strong>find the chunks that are most relevant to the question</strong>.</p>
<p>Let's take our previous example.</p>
<p>The user asks:</p>
<blockquote>
<p>"How many days can I work from home?"</p>
</blockquote>
<p>We convert the question into an embedding and compare it with the embeddings of our stored chunks.</p>
<p>Conceptually, we might get something like:</p>
<pre><code class="language-text">Question
   ↓
Question Embedding
   ↓
Compare with chunk embeddings
   ↓
┌─────────────────────────────────────┐
│ Chunk 1 → 0.91                       │
│ Chunk 2 → 0.32                       │
│ Chunk 3 → 0.18                       │
└─────────────────────────────────────┘
</code></pre>
<p>Here, the numbers represent how similar each chunk is to the question.</p>
<p><strong>Chunk 1</strong> has the highest similarity, so it is the most relevant chunk.</p>
<p>We can then retrieve the top relevant chunks and provide them as context to the LLM.</p>
<pre><code class="language-text">User Question
      ↓
   Embedding
      ↓
Semantic Search
      ↓
Relevant Chunks
      ↓
Question + Retrieved Context
      ↓
     LLM
      ↓
    Answer
</code></pre>
<p>This is the core retrieval step in RAG.</p>
<p>Instead of sending the entire document to the LLM, we first search through our stored knowledge and retrieve only the information that is relevant to the current question.</p>
<p>There are different ways to perform this similarity search. One common approach is <strong>cosine similarity</strong>, which measures how close two embedding vectors are. I implemented this calculation directly in Python instead of hiding it behind a vector database, so I could understand the retrieval process better.</p>
<p>In production systems, we usually don't calculate this against every chunk ourselves. A <strong>vector database</strong> can store the embeddings and perform these similarity searches efficiently.</p>
<p>So now we have most of the core pieces:</p>
<p><strong>Document → Chunking → Embeddings → Retrieval → Relevant Context</strong></p>
<p>But we still need the final step.</p>
<p><strong>How do we take this retrieved context and actually get an answer from the LLM?</strong></p>
<h3>Generation</h3>
<p>So, we now have the relevant context.</p>
<p>But how do we actually use it to get an answer from the LLM?</p>
<p>The answer is simple: <strong>we pass the retrieved context along with the user's question to the LLM.</strong></p>
<p>For example, the user asks:</p>
<blockquote>
<p>"How many days can I work from home?"</p>
</blockquote>
<p>Our retrieval step found this relevant chunk:</p>
<blockquote>
<p>"Employees can work remotely for up to 3 days per week."</p>
</blockquote>
<p>We can now provide both pieces of information to the LLM:</p>
<pre><code class="language-text">Context:
Employees can work remotely for up to 3 days per week.

Question:
How many days can I work from home?
</code></pre>
<p>The LLM uses the provided context to generate the answer:</p>
<blockquote>
<p>"You can work remotely for up to 3 days per week."</p>
</blockquote>
<p>And that's the <strong>Generation</strong> part of RAG.</p>
<p>The LLM isn't responsible for searching through the entire document. The retrieval system has already found the relevant information.</p>
<p>The LLM's job is now to <strong>use that context and generate a natural-language response</strong> to the user's question.</p>
<p>So if we put everything together:</p>
<p><strong>Document → Chunking → Embeddings → Retrieval → Context → LLM → Answer</strong></p>
<p>And that's the basic RAG pipeline.</p>
<hr />
<h2>How I Implemented the RAG Pipeline</h2>
<p>Now that we understand the basic pieces of RAG, let's look at how I used them in this project.</p>
<p>The implementation has two main flows:</p>
<ol>
<li><p><strong>Upload flow</strong> — turn a Markdown document into searchable chunks.</p>
</li>
<li><p><strong>Query flow</strong> — take a user's question, retrieve useful context, and generate an answer.</p>
</li>
</ol>
<hr />
<h3>Tech Stack</h3>
<ul>
<li><p><strong>Backend:</strong> Python and FastAPI</p>
</li>
<li><p><strong>Database:</strong> PostgreSQL with SQLAlchemy</p>
</li>
<li><p><strong>AI:</strong> OpenAI-compatible text and embedding APIs, with Ollama for local generation</p>
</li>
<li><p><strong>Retrieval:</strong> NumPy cosine similarity</p>
</li>
<li><p><strong>Tooling:</strong> Docker and uv</p>
</li>
</ul>
<hr />
<h3>Upload Flow</h3>
<p>The upload flow starts when a user uploads a <strong>Markdown file</strong> with a short description of what the document contains.</p>
<p>First, the document is parsed into a <strong>structured tree</strong> containing its title, sections, subsections, headings, and content. I use the main LLM for this parsing step because Markdown files can have different structures. If the request times out, encounters a network error, or returns invalid JSON, a simpler <strong>regex-based parser</strong> falls back to reading the Markdown headings.</p>
<p>Next, the section tree is converted into chunks. Instead of cutting the document at arbitrary positions, each section becomes a <strong>meaningful chunk</strong> and keeps a breadcrumb such as:</p>
<pre><code class="language-text">Troubleshooting &gt; Files Not Syncing
</code></pre>
<p>If a section is larger than <strong>2,000 characters</strong>, it is split into smaller windows with a <strong>200-character overlap</strong>. The overlap prevents useful context near a boundary from being lost.</p>
<p>Each chunk is then sent to the configured <strong>embedding API</strong>. The returned embedding, chunk text, heading information, and breadcrumb are stored in <strong>PostgreSQL</strong>. Embeddings are currently stored as JSON arrays rather than using pgvector directly.</p>
<p>I chose this approach because it keeps the implementation easy to inspect. PostgreSQL stores all document data in one place, while calculating <strong>cosine similarity in Python</strong> makes the retrieval logic visible instead of hiding it behind a vector database. This works well for a learning project and a small number of documents, although pgvector would be a better option at a larger scale.</p>
<p><img src="https://raw.githubusercontent.com/Prathamesh017/universal-vault/main/apps/images/upload-flow.png" alt="Upload Flow Diagram" /></p>
<hr />
<h3>Query Flow</h3>
<p>The query flow begins by understanding the question, then decides how much context is needed before searching the document.</p>
<h4>Step 1 — Handle Follow-Up Questions</h4>
<p>The system checks for phrases such as <em>"tell me more"</em>, <em>"continue"</em>, and <em>"what about"</em>. A normal question continues unchanged. For a follow-up, the last <strong>five cached question-and-answer turns</strong> for the document are loaded.</p>
<p>The main model, with Ollama as fallback, then chooses one of three actions:</p>
<ul>
<li><p><strong>Answer:</strong> the recent history already contains enough context, so return an answer immediately.</p>
</li>
<li><p><strong>Clarify:</strong> the reference is unclear, so ask the user what topic they meant.</p>
</li>
<li><p><strong>Search:</strong> rewrite the follow-up as a standalone question and continue through the pipeline.</p>
</li>
</ul>
<p>If no history exists, the API asks the user to clarify rather than searching for an ambiguous phrase.</p>
<h4>Step 2 — Classify the Question</h4>
<p>The prepared question is classified using simple keyword rules:</p>
<ul>
<li><p><strong>Definition:</strong> <code>what is</code>, <code>define</code>, <code>explain</code>, or <code>what does</code></p>
</li>
<li><p><strong>How-to:</strong> <code>how do I</code>, <code>how to</code>, <code>steps</code>, <code>install</code>, or <code>set up</code></p>
</li>
<li><p><strong>Comparison:</strong> <code>vs</code>, <code>compared to</code>, <code>difference</code>, or <code>better</code></p>
</li>
<li><p><strong>Troubleshooting:</strong> <code>why doesn't</code>, <code>error</code>, <code>problem</code>, <code>fix</code>, or <code>broken</code></p>
</li>
<li><p><strong>Factual:</strong> the default when no other rule matches</p>
</li>
</ul>
<p>The type then chooses the retrieval mode:</p>
<ul>
<li><p><strong>Quick:</strong> definition and factual — top 3 chunks, threshold 0.78</p>
</li>
<li><p><strong>Balanced:</strong> how-to — top 5 chunks, threshold 0.68</p>
</li>
<li><p><strong>Detailed:</strong> comparison and troubleshooting — top 10 chunks, threshold 0.58</p>
</li>
</ul>
<h4>Step 3 — Check the Cache in Two Ways</h4>
<p>Before validation or retrieval, the system checks whether the question has already been answered:</p>
<ol>
<li><p><strong>Exact match:</strong> normalize the question to lowercase and collapse extra spaces.</p>
</li>
<li><p><strong>Semantic match:</strong> embed the question and compare it with previous question embeddings. A score of <strong>0.90 or higher</strong> is a cache hit.</p>
</li>
</ol>
<p>On either hit, the stored answer is returned immediately with <code>cached: true</code>. Scores between 0.85 and 0.90 are logged as near misses but continue through the pipeline.</p>
<h4>Step 4 — Validate the Question</h4>
<p>On a cache miss, the main model checks that the question is both <strong>related to the document description</strong> and written as a <strong>clear question rather than disconnected keywords</strong>. Invalid questions return <code>isValid: false</code> with a short reason. If the validation call itself fails, retrieval is allowed to continue instead of blocking the user.</p>
<h4>Step 5 — Retrieve Relevant Chunks</h4>
<p>The question is converted into an embedding and compared with every stored chunk embedding using <strong>cosine similarity in Python</strong>. Chunks must:</p>
<ul>
<li><p>pass the selected mode's minimum threshold,</p>
</li>
<li><p>remain within <strong>0.12</strong> of the best score,</p>
</li>
<li><p>and fit within the mode's <code>top_k</code> limit.</p>
</li>
</ul>
<p>If nothing passes the threshold but weak matches exist, those nearest chunks are given to the main model to rewrite the query. Ollama is used if the main model fails. Retrieval is then attempted <strong>one more time</strong>. If the retry also fails, answer generation receives no chunks and returns no match.</p>
<h4>Step 6 — Generate the Answer</h4>
<p>The retrieved chunks and prepared question are sent to an answer model. A deterministic keyword check chooses the model order:</p>
<ul>
<li><p><strong>Simple question:</strong> Ollama first, then the main model as fallback.</p>
</li>
<li><p><strong>Other question:</strong> main model first, then Ollama as fallback.</p>
</li>
</ul>
<p>The model returns structured JSON containing whether the chunks actually satisfy the question and the answer message. If the chunks do not support an answer, or both models fail, the API returns <strong>"No matching information was found for your question."</strong></p>
<h4>Step 7 — Save and Return</h4>
<p>Only successful answers are saved to <code>conversation_history</code>, together with the normalized question and its embedding. This row is used by both future cache lookups and follow-up resolution.</p>
<p>The result is logged as <code>pipeline.answer_found</code> or <code>pipeline.answer_not_found</code>, and the final response includes the original question, the query actually used for retrieval, whether retrieval rewrote it, its question type, retrieval mode, and cache status.</p>
<p><img src="https://raw.githubusercontent.com/Prathamesh017/universal-vault/main/apps/images/User%20Question%20Validation-2026-09-29-071622.png" alt="Query Flow" /></p>
<p>The <strong>cache</strong> is important for cost and latency, while <strong>validation and answer checking</strong> are important for correctness. <strong>Query rewriting</strong> improves recall when the user's wording differs from the document. Together, these steps make the system more reliable than a pipeline that simply retrieves the closest chunk and always asks an LLM to answer.</p>
<hr />
<h3>Key Features Beyond Basic RAG</h3>
<p>The core RAG flow is only part of the implementation. I added several features around it to make the system more useful and resilient.</p>
<h4>Conversation History</h4>
<p>Conversation history makes short follow-ups understandable. Instead of embedding <em>"tell me more"</em> directly, the system uses recent turns to determine what the user means.</p>
<h4>Query Rewriting</h4>
<p>Query rewriting gives retrieval one more chance when the original question does not find a sufficiently similar chunk. The nearest weak chunks provide useful vocabulary for producing a better standalone query.</p>
<h4>Model Routing</h4>
<p>Model routing uses a deterministic keyword check. Simple questions such as definitions can be handled by a smaller local Ollama model, while questions involving comparison, architecture, integration, or troubleshooting go to the main model.</p>
<h4>Fallbacks</h4>
<p>Fallbacks prevent one model failure from breaking the whole request. Answer generation, follow-up resolution, and query rewriting can fall back to Ollama if the main model fails. Simple answer generation can also fall back in the opposite direction.</p>
<h4>Caching</h4>
<p>Caching supports both exact and semantic matches. Found answers are stored with the normalized question and its embedding, so repeated and closely paraphrased questions can return immediately.</p>
<h4>Metrics and Rate Limiting</h4>
<p>The project also records query events for metrics and applies simple per-IP rate limits. These are not part of RAG itself, but they helped me understand the operational pieces that surround a real question-answering system.</p>
<hr />
<h2>Limitations</h2>
<p>This is a <strong>learning and portfolio project</strong> focused on understanding how a RAG system works end to end, so some production concerns are intentionally out of scope.</p>
<ul>
<li><p><strong>No frontend yet.</strong> The main goal was the RAG pipeline itself; the API is used through Swagger (<code>/docs</code>) or <code>curl</code>.</p>
</li>
<li><p><strong>Markdown only.</strong> Chunking relies on Markdown headings. Other formats such as PDF, DOCX, and HTML could be supported by converting them to Markdown before upload.</p>
</li>
<li><p><strong>pgvector is not used directly yet.</strong> Embeddings are stored as JSON arrays in PostgreSQL and scored with cosine similarity in Python. This is fine for a few documents, but pgvector with an index would move the search into PostgreSQL and scale much better.</p>
</li>
<li><p><strong>Embedding search only.</strong> There is no keyword or full-text search alongside it, so exact terms such as error codes and version numbers can be missed.</p>
</li>
<li><p><strong>One document per query.</strong> Questions cannot currently span multiple documents.</p>
</li>
<li><p><strong>Shared history per document.</strong> Follow-ups use the last five question-and-answer turns for the document, rather than maintaining separate history per user or session.</p>
</li>
</ul>
<hr />
<h2>Conclusion</h2>
<p>Building this project showed me that RAG is much more than placing a document inside an LLM prompt. The quality of the final answer depends on <strong>how the document is chunked, how context is retrieved, how weak queries are rewritten, and how unsupported answers are rejected</strong>.</p>
<p>The most useful lesson was that many reliability improvements happen <strong>around the LLM</strong>, not inside it. Caching reduces repeated work, validation avoids irrelevant searches, conversation history makes follow-ups meaningful, and model routing balances local generation with a stronger fallback.</p>
<p>This implementation is intentionally small and transparent, but it covers the complete path from <strong>document upload to a grounded answer</strong>.</p>
]]></content:encoded></item><item><title><![CDATA[Building a Temporal-Based CI Orchestrator in Go]]></title><description><![CDATA[Introduction
You push code to GitHub and your CI pipeline starts running. Then it happens:

Lint fails.

You fix it and push again.

Tests fail.

You fix that, push again, and now the build fails.Afte]]></description><link>https://prathamesh-thakare.hashnode.dev/building-a-temporal-based-ci-orchestrator-in-go</link><guid isPermaLink="true">https://prathamesh-thakare.hashnode.dev/building-a-temporal-based-ci-orchestrator-in-go</guid><category><![CDATA[temporal]]></category><category><![CDATA[Orchestration]]></category><category><![CDATA[ci-cd]]></category><category><![CDATA[golang]]></category><category><![CDATA[Bubbletea]]></category><dc:creator><![CDATA[Prathamesh Thakare]]></dc:creator><pubDate>Wed, 16 Sep 2026 11:51:59 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/648c30cf85f011423c0e5c2a/fd77b05c-91e3-4ec3-850d-d61d62e5b008.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Introduction</h4>
<p>You push code to GitHub and your CI pipeline starts running. Then it happens:</p>
<blockquote>
<p>Lint fails.</p>
</blockquote>
<p>You fix it and push again.</p>
<blockquote>
<p>Tests fail.</p>
</blockquote>
<p>You fix that, push again, and now the build fails.After a few iterations, you've spent more time waiting for CI than actually fixing the problem.</p>
<p>The frustrating part isn't that the pipeline failed. It's that every failure usually means <strong>starting the entire pipeline again</strong>, even when only one step was the problem.</p>
<p>Your first instinct is to automate it with a shell script:</p>
<pre><code class="language-bash">#!/bin/bash

go lint ./...
go test ./...
go build ./...
./scripts/deploy.sh
</code></pre>
<p>It works. The commands run one after another, and the script exits as soon as something fails.</p>
<p>But shell scripts have limits:</p>
<ul>
<li>A failure means starting from the beginning.</li>
<li>Retrying a single step usually means rerunning the script manually.</li>
<li>If the machine crashes halfway through, the execution state is gone.</li>
<li>There's no built-in way to pause for approvals or inspect progress while it's running.</li>
</ul>
<p>This is the problem <strong>Conductor</strong> tries to solve.</p>
<p>Instead of writing imperative scripts, you describe your CI pipeline in a simple YAML file, and <strong>Temporal</strong> executes it as a durable workflow.</p>
<p>That changes how a pipeline behaves.</p>
<hr />
<h4>Tech Stack</h4>
<p>Conductor is built around a few simple technologies:</p>
<ul>
<li><strong>Go</strong> — The main language used to build Conductor.</li>
<li><strong>Temporal</strong> — Handles workflow orchestration, retries, state, and durable execution.</li>
<li><strong>Bubble Tea</strong> — Powers the terminal UI (TUI).</li>
<li><strong>Docker Compose</strong> — Used to run Temporal and its dependencies locally.</li>
<li><strong>PostgreSQL</strong> — Used by the local Temporal setup for persistence.</li>
</ul>
<p>For the terminal UI, I'm using <strong>Bubble Tea</strong> with the <strong>Dracula theme</strong> to keep the interface simple and consistent.</p>
<p>The local development environment can be started with Docker Compose, which sets up the Temporal server, Temporal UI, and PostgreSQL.</p>
<p>So the overall stack looks like this:</p>
<pre><code class="language-text">                    Conductor
                       │
             ┌─────────┴─────────┐
             │                   │
        Bubble Tea           Temporal
          (TUI)             (Workflow Engine)
                                 │
                         ┌───────┴───────┐
                         │               │
                     PostgreSQL      Temporal UI
                         │
                    Docker Compose
</code></pre>
<hr />
<h4>What is Conductor CI</h4>
<p><strong>Conductor turns a YAML pipeline into a Temporal workflow that you can monitor and control from your terminal.</strong></p>
<p>The idea is simple:</p>
<ul>
<li><strong>YAML</strong> defines what the pipeline should do.</li>
<li><strong>Temporal</strong> executes and manages the workflow.</li>
<li><strong>Terminal UI</strong> lets you monitor and control it.</li>
</ul>
<p>Instead of writing a shell script, you describe your pipeline in YAML:</p>
<pre><code class="language-yaml">steps:
  - name: lint
    run: go lint ./...

  - name: test
    run: go test ./...

  - name: build
    run: go build ./...

  - name: deploy
    run: ./scripts/deploy.sh
</code></pre>
<p>Conductor takes this definition and runs it as a Temporal workflow. This gives your pipeline features that a simple script doesn't have. Lets understand in details</p>
<hr />
<h4>User Journey</h4>
<p>Once you have a <code>workflow.yaml</code> in your project, getting started is simple:</p>
<pre><code class="language-bash">$ task run
</code></pre>
<p>or execute the go binary from your <code>workflow.yaml</code> location folder.
Checkout <a href="https://github.com/Prathamesh017/conductor-ci/blob/main/README.md">Readme </a> for more details</p>
<p>Conductor starts in the terminal and gives you two choices:</p>
<p><img src="https://github-production-user-asset-6210df.s3.amazonaws.com/92641320/652918445-a31c208b-2a76-47e8-bcdc-afa085e067c9.png?X-Amz-Algorithm=AWS4-HMAC-SHA256&amp;X-Amz-Credential=AKIAVCODYLSA53PQK4ZA%2F20260916%2Fus-east-1%2Fs3%2Faws4_request&amp;X-Amz-Date=20260916T120707Z&amp;X-Amz-Expires=300&amp;X-Amz-Signature=11f9871409b8d972e6b1322cff5a6a807c33c82c22931611f9f82a4dcdb0f0b3&amp;X-Amz-SignedHeaders=host&amp;response-content-type=image%2Fpng" alt="Screenshot 2026-09-16 at 1 28 24 PM" /></p>
<h5>Option 1: <strong>Validate Workflow</strong></h5>
<p>Before running anything, you can check whether your workflow.yaml is valid.</p>
<p>Select validate-workflow-yaml and Conductor parses and validates the YAML without connecting to Temporal or starting a workflow.</p>
<p><img src="https://github-production-user-asset-6210df.s3.amazonaws.com/92641320/652919450-021c5669-9f69-4871-bcb9-c9093f949b9d.png?X-Amz-Algorithm=AWS4-HMAC-SHA256&amp;X-Amz-Credential=AKIAVCODYLSA53PQK4ZA%2F20260916%2Fus-east-1%2Fs3%2Faws4_request&amp;X-Amz-Date=20260916T120957Z&amp;X-Amz-Expires=300&amp;X-Amz-Signature=1f74ab434a6aceff3a24db2bd2223701e6ee6d00cee4a0461ea32c34d38b3dad&amp;X-Amz-SignedHeaders=host&amp;response-content-type=image%2Fpng" alt="image2" /></p>
<p>The checks are grouped to make it easy to understand what went wrong.</p>
<p>Nothing has been started at this point. No Temporal server. No worker. No side effects. Just validation.</p>
<p>You fix your YAML and run the validation again.</p>
<hr />
<h5>Option 2: <strong>Start Workflow</strong></h5>
<p>Once you're ready to run the pipeline, select start-workflow.</p>
<p>Conductor first performs the same validation. This makes sure an invalid configuration never reaches Temporal. If validation passes, Conductor moves on to Temporal:</p>
<h5>The TUI: Watch It Happen in Real-Time</h5>
<p>Now TUI Takes Over .Instead of printing a continuous stream of logs, Conductor gives you a live view of the workflow:</p>
<img src="https://github.com/user-attachments/assets/e6a2cabf-f787-42b9-aefb-cde4f7963a2a" alt="SS" style="display:block;margin:0 auto" />

<p>Each task has a simple status:</p>
<p>✓ — Task passed ✗ — Task failed ⟳ — Task is running ⏳ — Task is queued</p>
<p>You can also see elapsed time, stage progress, and interact with the workflow when needed.</p>
<p>The goal is simple:</p>
<p>Instead of watching logs, you watch the workflow itself.</p>
<hr />
<h4>What You Can Do During a Workflow (Features)</h4>
<p>Now the workflow is running. You're not passively watching. You can control it.</p>
<h5><strong>Automatic retries</strong></h5>
<p>Tasks can retry on their own when they fail. That count lives on the task:</p>
<pre><code class="language-yaml">tasks:
  - name: test
    script: go test ./...
    retries: 2
</code></pre>
<p>If <code>test</code> fails, Temporal retries it according to that count. You do not press anything. The engine does it.</p>
<hr />
<h5><strong>Manual retry</strong></h5>
<p>If a task still fails — and you have fixed the problem — you can retry <strong>just that task</strong> from the TUI.</p>
<pre><code class="language-text">build
  ✗ build
</code></pre>
<p>Press <code>r</code> to retry:</p>
<pre><code class="language-text">build
  ⟳ build

build
  ✓ build
</code></pre>
<p>Completed tasks do not run again. The workflow continues from the failed task.</p>
<hr />
<h5><strong>Approval gates</strong></h5>
<p>Some stages should not continue until a person says yes:</p>
<pre><code class="language-yaml">execution:
  - stage: production
    tasks: [deploy]
    requires_approval: true
</code></pre>
<p>The workflow pauses at that stage:</p>
<pre><code class="language-text">production
  ⏳ deploy (awaiting approval...)
</code></pre>
<p>Press <code>a</code> to approve. Pressing <code>a</code> does not run deploy by itself. It lets the workflow continue. The deploy activity then goes onto the queue, and a worker picks it up.</p>
<hr />
<h5><strong>Cancel a workflow</strong></h5>
<p>If something goes wrong, you can cancel the run from the TUI.</p>
<pre><code class="language-text">Terminate workflow? (y/n): y

Workflow terminated by user
</code></pre>
<p>The workflow stops. Its execution history is preserved.</p>
<h5>Parallel and sequential tasks</h5>
<p>You control how tasks <strong>inside a stage</strong> execute:</p>
<pre><code class="language-yaml">execution:
  - stage: quality
    tasks: [lint, format]
    mode: parallel

  - stage: deploy
    tasks: [deploy-staging, health-check]
    mode: sequential
</code></pre>
<table>
<thead>
<tr>
<th>Mode</th>
<th>Behavior</th>
</tr>
</thead>
<tbody><tr>
<td><code>parallel</code></td>
<td>Tasks in the stage run at the same time.</td>
</tr>
<tr>
<td><code>sequential</code></td>
<td>The next task starts only after the previous one finishes.</td>
</tr>
</tbody></table>
<p>Stages themselves still run top to bottom. Parallel is inside a stage, not across the whole pipeline.</p>
<hr />
<h4>Lets Understand  Temporal</h4>
<p><a href="https://temporal.io/">Temporal</a> is a durable execution engine. You describe a long-running process as a <strong>workflow</strong> — a function that says what should happen, in what order, and what to do when something fails.</p>
<p>Temporal records every step of that function as <strong>history</strong>. If a machine dies, a deploy happens, or the process restarts, Temporal <strong>replays</strong> that history and continues. The run does not start over. It resumes.</p>
<p>That is the whole pitch, independent of CI. Temporal is not a dashboard for jobs. It is infrastructure for processes that must not get lost.</p>
<hr />
<h5>How Conductor Actually use Them</h5>
<p>These are all the pieces we use,</p>
<table>
<thead>
<tr>
<th>Piece</th>
<th>Role</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Workflow</strong></td>
<td>The process. What happens next.</td>
</tr>
<tr>
<td><strong>Activity</strong></td>
<td>A real-world step: charge a card, call an API, run a script.</td>
</tr>
<tr>
<td><strong>Signal</strong></td>
<td>The outside world writes <em>into</em> a running process.</td>
</tr>
<tr>
<td><strong>Query</strong></td>
<td>The outside world <em>reads</em> it without changing it.</td>
</tr>
</tbody></table>
<p>A CI pipeline is that pattern. Lint, test, build, wait for approval, deploy. The scripts are activities. The human at the keyboard is signals. The live status on screen is a query.</p>
<p>For us, that process starts as YAML in the repo — not as something you click together in Temporal:</p>
<pre><code class="language-yaml">name: PR Validation

tasks:
  - name: lint
    script: echo "Linting From Terminal..."
    retries: 0
  - name: test
    script: echo "Testing From Terminal..."
    retries: 2
  - name: build
    script: echo "Building From Terminal..."
    retries: 1
  - name: deploy
    script: echo "Deploying From Terminal..."
    retries: 0

execution:
  - stage: validation
    tasks: [lint, test]
    mode: sequential
  - stage: build
    tasks: [build]
    mode: sequential
  - stage: deploy
    tasks: [deploy]
    mode: sequential
    requires_approval: true
</code></pre>
<p>Temporal never sees this file. It sees a workflow that <em>means</em> this file: stages in order, each task as an activity, deploy paused until a human arrives.</p>
<hr />
<h5>The moving parts</h5>
<p>Temporal is a <strong>server</strong> plus <strong>workers</strong>. The server is the source of truth. Workers are how work actually happens.</p>
<p><strong>The server and the task queue</strong></p>
<p>When a workflow wants something done — “run this lint script,” “charge this card” — it does not run that code itself. It <strong>schedules an activity</strong>. That activity is placed on a <strong>task queue</strong>: a named inbox on the Temporal server.</p>
<p>A task queue is not a thread. It is a mailbox.</p>
<pre><code class="language-yaml">task_queue: conductor-ci-queue

# The workflow says:
#   put this activity on conductor-ci-queue
# The server holds it until a worker is listening.
</code></pre>
<hr />
<p><strong>What a worker is</strong></p>
<p>A <strong>worker</strong> is a process you run. It connects to the Temporal server, <strong>registers</strong> the workflow and activity functions it knows how to execute, and <strong>polls</strong> a task queue.</p>
<p>When a task appears on that queue, a worker picks it up and runs the matching function. When it finishes, the result goes back to the server, and the workflow continues.</p>
<p>If no worker is running, the workflow is still recorded. Nothing executes until a worker comes online and starts polling. That split is the point: the server remembers the process; the worker is replaceable.</p>
<p>Registration is how the worker announces itself:</p>
<pre><code class="language-yaml">worker:
  listens_on: conductor-ci-queue
  workflows:
    - prWorkflow          # the process
  activities:
    - runTaskActivity     # the real work (your scripts)
</code></pre>
<p>Until that happens, Temporal has history, but no hands.</p>
<hr />
<h5>How an activity actually runs</h5>
<p>The path is always <strong>queue → worker → result</strong>.</p>
<pre><code class="language-mermaid">flowchart TD
  A[Workflow decides: run lint] --&gt; B[Activity lands on the task queue]
  B --&gt; C[Worker polls and picks it up]
  C --&gt; D[Activity runs the script]
  D --&gt; E[Result returns to the workflow]
  E --&gt; F[Next stage, retry, or wait]
</code></pre>
<ol>
<li>The workflow reaches a step that needs the real world.</li>
<li>It schedules an <strong>activity</strong> onto the task queue.</li>
<li>A worker polling that queue <strong>picks the activity up</strong>.</li>
<li>The worker runs the activity — in our case, the task’s <code>script</code>, in the project directory.</li>
<li>Success or failure goes back to Temporal. The workflow decides what is next.</li>
</ol>
<p>Activities are allowed to fail. That is why they exist. The workflow stays deterministic; the activity is allowed to touch the world.</p>
<p>Timeouts and retries live on the activity. In YAML, that is the <code>retries</code> field on a task:</p>
<pre><code class="language-yaml">- name: test
  script: echo "Testing From Terminal..."
  retries: 2    # Temporal tries the activity up to 3 times on its own
</code></pre>
<hr />
<h5><strong>Temporal UI</strong></h5>
<p>The screen is the product. Temporal is the engine behind it.
It gives you a visual view of your workflow execution, making it easy to see:</p>
<ul>
<li>Which workflows are currently running or completed</li>
<li>The status of each workflow</li>
<li>Individual activities and their execution state</li>
<li>Workflow execution history</li>
<li>Retries and failures</li>
<li>Signals and other workflow events</li>
</ul>
<p>For Conductor, this is especially useful because you can see the pipeline you defined in YAML being executed by Temporal in real time.</p>
<p><img src="https://github-production-user-asset-6210df.s3.amazonaws.com/92641320/652954970-f4b7eab2-4264-4369-a08e-0294a49f10c6.png?X-Amz-Algorithm=AWS4-HMAC-SHA256&amp;X-Amz-Credential=AKIAVCODYLSA53PQK4ZA%2F20260916%2Fus-east-1%2Fs3%2Faws4_request&amp;X-Amz-Date=20260916T123045Z&amp;X-Amz-Expires=300&amp;X-Amz-Signature=af275792b55f28c8b632961e62cb0f0985625f83b45367ee5f5a31094e1789e9&amp;X-Amz-SignedHeaders=host&amp;response-content-type=image%2Fpng" alt="Temporal UI" /></p>
<hr />
<h5>Signals</h5>
<p>A <strong>signal</strong> is how the outside world <strong>writes</strong> into a running workflow. It has a name. It becomes part of history. If the worker restarts, Temporal still knows the signal happened — or that the workflow is still waiting for it.</p>
<p>This is not a sleep. It is not a prompt that vanishes when you close the terminal. The workflow <strong>blocks on a named event</strong>. Until that event arrives, the process is paused and durable.</p>
<p>In general, signals are how you inject a human or another system: payment received, customer cancelled, manager approved.</p>
<p>We use two. Both are born from the YAML, then sent from the UI:</p>
<pre><code class="language-yaml"># In the file: this stage must not start until a human says yes
- stage: deploy
  tasks: [deploy]
  mode: sequential
  requires_approval: true
</code></pre>
<table>
<thead>
<tr>
<th>Signal</th>
<th>Key</th>
<th>What it means</th>
</tr>
</thead>
<tbody><tr>
<td><code>approve</code></td>
<td><code>a</code></td>
<td>The gated stage may continue.</td>
</tr>
<tr>
<td><code>retry</code></td>
<td><code>r</code></td>
<td>Run the failed task again.</td>
</tr>
</tbody></table>
<hr />
<h5>Queries</h5>
<p>A <strong>query</strong> is how the outside world <strong>reads</strong> a running workflow. It does not change history. It does not wait for a step to finish. It asks: <em>what is your state right now?</em></p>
<p>In general, queries are how you build a UI on a process you do not own: order status, where this onboarding is, whether this is still waiting on approval.</p>
<p>We register one:</p>
<pre><code class="language-yaml">query: execution-state

# What the UI asks, on a short interval:
#   which stages are queued, running, passed, failed, or awaiting?
#   how long did each task take?
</code></pre>
<p>The UI is not inside the workflow. The workflow is running on the worker. So the terminal asks that query and paints whatever comes back. That is why the screen can show a task as running while the worker is still executing the activity, and why approval shows up as a pause instead of a frozen app.</p>
<hr />
<h4>Conclusion</h4>
<p>Building Conductor started as a way to explore how Temporal could be used for something familiar: CI pipelines.</p>
<p>Conductor is still a learning project, but building it gave me a much better understanding of Temporal, durable execution, signals, queries, and workflow orchestration.</p>
<p>You can find the complete project here:</p>
<p><a href="https://github.com/Prathamesh017/conductor-ci">Conductor CI on GitHub</a></p>
<p>Thanks for reading! 🚀</p>
]]></content:encoded></item><item><title><![CDATA[Understanding LSM Trees From First Principles




]]></title><description><![CDATA[Introduction
Before Starting on LSM Trees, Let's Rewind a Bit
Imagine it's 1970-something, and you're building a simple database for a manga library.
You need to store manga records:



manga_id
title]]></description><link>https://prathamesh-thakare.hashnode.dev/understanding-lsm-trees-from-first-principles</link><guid isPermaLink="true">https://prathamesh-thakare.hashnode.dev/understanding-lsm-trees-from-first-principles</guid><dc:creator><![CDATA[Prathamesh Thakare]]></dc:creator><pubDate>Fri, 04 Sep 2026 04:06:23 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/648c30cf85f011423c0e5c2a/aaa84758-8f36-4074-9ff8-663bef44ea35.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Introduction</h2>
<p>Before Starting on LSM Trees, Let's Rewind a Bit</p>
<p>Imagine it's <strong>1970-something</strong>, and you're building a simple database for a manga library.</p>
<p>You need to store manga records:</p>
<table>
<thead>
<tr>
<th>manga_id</th>
<th>title</th>
<th>author</th>
<th>chapters</th>
</tr>
</thead>
<tbody><tr>
<td>1</td>
<td>One Piece</td>
<td>Eiichiro Oda</td>
<td>1150+</td>
</tr>
<tr>
<td>2</td>
<td>Naruto</td>
<td>Masashi Kishimoto</td>
<td>700</td>
</tr>
<tr>
<td>3</td>
<td>Berserk</td>
<td>Kentaro Miura</td>
<td>370+</td>
</tr>
<tr>
<td>4</td>
<td>Death Note</td>
<td>Tsugumi Ohba</td>
<td>108</td>
</tr>
<tr>
<td>5</td>
<td>Dragon Ball</td>
<td>Akira Toriyama</td>
<td>519</td>
</tr>
</tbody></table>
<p>Your database lives on <strong>disk</strong>.If Someone asks:</p>
<blockquote>
<p>"Give me the manga with ID 3."</p>
</blockquote>
<p>You need to find that record as quickly as possible.So let's forget about fancy database data structures for a moment.</p>
<p><strong>How would you store these records if you had to build this yourself?</strong></p>
<h4>The Simplest Thing You Could Do</h4>
<p>The most obvious solution is to just write the records to disk one after another, in whatever order they arrive.</p>
<pre><code class="language-text">Disk:

Position 1 → Manga ID 5
Position 2 → Manga ID 2
Position 3 → Manga ID 8
Position 4 → Manga ID 1
Position 5 → Manga ID 7
</code></pre>
<p>Now someone asks:</p>
<blockquote>
<p>"Find me Manga ID 8."</p>
</blockquote>
<p>Well... we don't know where it is. So we start reading from the beginning:</p>
<pre><code class="language-text">Position 1 → Manga 5 → not it
Position 2 → Manga 2 → not it
Position 3 → Manga 8 → found it!
</code></pre>
<p>For five records, who cares? But what if we have <strong>a million manga</strong>?</p>
<p>Now finding one manga could mean reading hundreds of thousands of records before we get to the one we want.
And there's a problem with that:</p>
<blockquote>
<p><strong>Disk reads are expensive.</strong></p>
</blockquote>
<p>We don't want our database spending most of its time reading records that have nothing to do with the query. So maybe we can organize the data a little better.</p>
<h4>What If We Sort It?**</h4>
<p>Instead of writing manga in whatever order they arrive, we could keep them sorted by <code>manga_id</code>.</p>
<pre><code class="language-text">Disk:

Position 1 → Manga ID 1
Position 2 → Manga ID 2
Position 3 → Manga ID 5
Position 4 → Manga ID 7
Position 5 → Manga ID 8
</code></pre>
<p>That's already better.</p>
<p>If we're looking for Manga ID <code>8</code>, we know the IDs are ordered. Once we reach an ID greater than <code>8</code>, we can stop.</p>
<p>But there's still a problem.</p>
<p>With <strong>a million records</strong>, we might still have to read a huge number of records before getting to the one we want.
Sorting tells us <strong>how the data is ordered</strong>.</p>
<p>It doesn't tell us <strong>where to jump</strong>.</p>
<h4>What If We Had an Index?</h4>
<p>What if we kept a small amount of information separately that told us roughly where different ranges of IDs live?</p>
<p>Something like:</p>
<pre><code class="language-text">Index:

Manga IDs 1–250,000
        ↓
Disk Position 1

Manga IDs 250,001–500,000
        ↓
Disk Position 500,001

Manga IDs 500,001–750,000
        ↓
Disk Position 1,000,001
</code></pre>
<p>Now suppose we're looking for:</p>
<pre><code class="language-text">Manga ID 8
</code></pre>
<p>Instead of starting from the beginning and blindly scanning through records, we first check the index.</p>
<p>The index tells us:</p>
<pre><code class="language-text">Manga ID 8
    ↓
IDs 1–250,000
    ↓
Start around Disk Position 1
</code></pre>
<p>We have narrowed down the search before touching the actual records. And that's a pretty useful idea.</p>
<p>Instead of storing only the data, we also keep <strong>some information about where the data is</strong>.</p>
<p>The question now becomes:</p>
<blockquote>
<p><strong>How do we build an index that lets us find things quickly without the index itself becoming huge?</strong></p>
</blockquote>
<p>That's where <strong>B-Trees</strong> come in.</p>
<hr />
<h1>So, What Does A B-Tree Look Like?</h1>
<p>We now have a basic idea of what an index can do.</p>
<p>Instead of searching the entire dataset, we keep some information that helps us quickly figure out <strong>where to look</strong>.</p>
<p>But there's a new problem.</p>
<p>If we have a million, or even a billion, records, we can't just keep a million-entry index and scan through that index too.</p>
<p>We need the <strong>index itself to be searchable</strong>.</p>
<p>This is where the <strong>B-Tree</strong> comes in.</p>
<p>A B-Tree is essentially a tree designed specifically for storing and searching data efficiently on disk. Let's build a very small one.</p>
<h4>A Simple B-Tree</h4>
<p>Suppose our manga IDs are:</p>
<pre><code class="language-text">1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12
</code></pre>
<p>A simplified B-Tree might look something like this:</p>
<pre><code class="language-text">                         [ 4 | 8 ]
                       /     |     \
                      /      |      \
             [1 | 2 | 3] [5 | 6 | 7] [9 | 10 | 11 | 12]
</code></pre>
<p>The important thing to notice is that the tree doesn't have just <strong>one value per node</strong>.</p>
<p>A node can contain multiple keys.</p>
<p>The keys divide the data into ranges.</p>
<p>For the root:</p>
<pre><code class="language-text">                 [ 4 | 8 ]
</code></pre>
<p>we can think of it as three ranges:</p>
<pre><code class="language-text">        &lt; 4        4–8        &gt; 8
         ↓          ↓           ↓
    left child   middle child  right child
</code></pre>
<p>So the root doesn't contain every manga.</p>
<p>It simply helps us decide <strong>which part of the tree we need to look at</strong>.</p>
<h4>Searching for a Manga</h4>
<p>Let's say someone asks:</p>
<blockquote>
<p>"Find Manga ID 10."</p>
</blockquote>
<p>We start at the root:</p>
<pre><code class="language-text">[ 4 | 8 ]
</code></pre>
<p>We compare <code>10</code> with the keys:</p>
<pre><code class="language-text">10 &gt; 4
10 &gt; 8
</code></pre>
<p>So we know the answer, if it exists, must be in the <strong>right child</strong>.</p>
<p>We follow that pointer:</p>
<pre><code class="language-text">                         [ 4 | 8 ]
                              |
                              ↓
                    [9 | 10 | 11 | 12]
</code></pre>
<p>Now we search inside that node:</p>
<pre><code class="language-text">[9 | 10 | 11 | 12]
       ↑
     found!
</code></pre>
<p>That's it. We didn't scan all 12 records. We followed the path that could actually contain ID <code>10</code>. With a much larger tree, the same idea continues.</p>
<p>For example:</p>
<pre><code class="language-text">                         [ 40 | 80 ]
                       /      |      \
                      /       |       \
             [10|20|30] [50|60|70] [90|100|110]
</code></pre>
<p>If we're searching for <code>70</code>:</p>
<pre><code class="language-text">[40 | 80]
     |
     ↓
[50 | 60 | 70]
          |
        found
</code></pre>
<p>The tree keeps reducing the amount of data we need to examine.</p>
<h4>Why Does a B-Tree Have Multiple Keys Per Node?</h4>
<p>This is an important detail.</p>
<p>Remember where our data lives:</p>
<p><strong>On disk.</strong></p>
<p>Reading from disk is expensive, so we don't want a tree where every tiny node requires another disk read.</p>
<p>Instead, a B-Tree packs <strong>many keys and pointers into a single node</strong>, typically sized to work well with disk pages.</p>
<p>Conceptually:</p>
<pre><code class="language-text">Disk Page
┌──────────────────────────────────────┐
│ Key │ Key │ Key │ Key │ Key │ ...    │
│     │     │     │     │     │        │
│ Pointer │ Pointer │ Pointer │ ...     │
└──────────────────────────────────────┘
</code></pre>
<p>One disk read can therefore give us a lot of information.</p>
<p>The tree can have a very large number of records while remaining relatively shallow.</p>
<p>That's one of the main reasons B-Trees work so well for disk-based databases.</p>
<hr />
<h4>B-Trees In Production</h4>
<p>B-Tree and B+Tree variants are widely used in real database systems.</p>
<p>For example:</p>
<ul>
<li><strong>PostgreSQL</strong> uses B-tree indexes as its default index type.</li>
<li><strong>MySQL/InnoDB</strong> uses B+Tree-style indexes for its primary and secondary indexes.</li>
<li><strong>SQLite</strong> uses B-trees for tables and indexes.</li>
</ul>
<p>So when you create something like:</p>
<pre><code class="language-sql">CREATE INDEX idx_manga_id
ON manga(manga_id);
</code></pre>
<p>the database isn't necessarily doing some magical lookup.</p>
<p>It's maintaining a data structure that lets it efficiently navigate to the records — commonly a B-tree variant.</p>
<h4>So What's the Problem?</h4>
<p>At this point, B-Trees sound pretty great.</p>
<p>They give us:</p>
<ul>
<li>Fast lookups</li>
<li>Ordered data</li>
<li>Efficient range queries</li>
<li>An index that doesn't require scanning everything</li>
<li>Good performance with data stored on disk</li>
</ul>
<p>So why would we ever need anything else?</p>
<p>The answer becomes obvious when we stop <strong>reading</strong> and start <strong>writing</strong>.</p>
<p>Imagine our database already contains:</p>
<pre><code class="language-text">                         [ 4 | 8 ]
                       /     |     \
                      /      |      \
                 [1|2|3]  [5|6|7]  [9|10|11]
</code></pre>
<p>Now a new manga arrives:</p>
<pre><code class="language-text">Manga ID = 6
</code></pre>
<p>Where does it go?</p>
<p>It belongs inside:</p>
<pre><code class="language-text">[5 | 6 | 7]
</code></pre>
<p>That sounds easy.But real database pages have limited space.
Suppose the node is already full:</p>
<pre><code class="language-text">[5 | 6 | 7 | 8]
</code></pre>
<p>and we need to insert:</p>
<pre><code class="language-text">6
</code></pre>
<p>The database can't simply keep adding data forever.It may need to <strong>split the node</strong>.
Something like:</p>
<pre><code class="language-text">Before:

              [ 4 | 8 ]
             /
      [5 | 6 | 7 | 8]
</code></pre>
<p>After the split, the tree structure may need to change:</p>
<pre><code class="language-text">              [ 4 | 6 | 8 ]
             /     |     \
            /      |      \
       [5]        [7]     [...]
</code></pre>
<p>The exact details are more complicated in a real B-Tree/B+Tree, but the important part is this:</p>
<blockquote>
<p><strong>A write isn't necessarily just writing the new record.</strong></p>
</blockquote>
<p>The database may need to:</p>
<ol>
<li>Find the correct page.</li>
<li>Read that page.</li>
<li>Modify it.</li>
<li>Potentially split the page.</li>
<li>Update parent nodes.</li>
<li>Potentially split those parents too.</li>
<li>Write multiple modified pages back to disk.</li>
</ol>
<p>And remember:</p>
<p><strong>Disk I/O is expensive.</strong></p>
<h4>Now Imagine Heavy Writes</h4>
<p>Suppose we're building something where data is arriving constantly:</p>
<pre><code class="language-text">10,000 writes/sec
100,000 writes/sec
1,000,000 writes/sec
</code></pre>
<p>Every write potentially involves modifying existing structures that are already sitting on disk.
The problem isn't that B-Trees are bad. They are actually extremely good at what they were designed to do.</p>
<p>The problem is that <strong>frequent random writes to disk can become expensive</strong>.
And this gives us a new question:</p>
<blockquote>
<p><strong>What if we didn't immediately modify the existing data structure on disk every time a write arrived?</strong></p>
</blockquote>
<p>What if we could accept writes somewhere much faster first, and deal with organizing them later?</p>
<p>That question takes us in a completely different direction.</p>
<p>And that's where <strong>LSM Trees</strong> begin.</p>
<hr />
<h2>LSM Tree</h2>
<p>We just saw the main limitation of a B-Tree:</p>
<blockquote>
<p><strong>Writes modify existing data on disk.</strong></p>
</blockquote>
<p>What if we avoided that?</p>
<p>Instead of immediately finding the right place on disk and modifying it, we can <strong>keep new writes in memory and write them to disk sequentially in batches</strong>.</p>
<p>That's the basic idea behind an <strong>LSM Tree (Log-Structured Merge Tree)</strong>.</p>
<p>A write follows roughly this path:</p>
<pre><code class="language-text">User Write
    ↓
  WAL
    ↓
Memtable
    ↓
  Flush
    ↓
SSTable
    ↓
Compaction
</code></pre>
<h3>Lets Understand All Components in Details</h3>
<h4>WAL (Write Ahead Logs)</h4>
<p>The first problem we need to solve is <strong>durability</strong>.
Our writes are going into the Memtable, which lives in memory. That's fast, but memory is not permanent.</p>
<p>Imagine we receive:</p>
<pre><code class="language-text">PUT manga:10 Berserk
</code></pre>
<p>If we put it directly into the Memtable:</p>
<pre><code class="language-text">Memtable

manga:10 → Berserk
</code></pre>
<p>and the process crashes, the write is gone.
So we first record the operation on disk in a <strong>Write-Ahead Log (WAL)</strong> and then put it into the Memtable:</p>
<pre><code class="language-text">User Write
    ↓
   WAL
    ↓
Memtable
</code></pre>
<p>A WAL is a common durability mechanism used by many databases, including traditional B-Tree-based databases. The idea is simply to <strong>record the operation before modifying the actual database state</strong>.</p>
<p>The WAL is an <strong>append-only log</strong>:</p>
<pre><code class="language-text">PUT manga:10 Berserk
PUT manga:11 One Piece
DELETE manga:12
</code></pre>
<p>We don't modify old entries. New operations are appended to the end.</p>
<p>The order is important:</p>
<pre><code class="language-text">1. Write to WAL
2. Write to Memtable
</code></pre>
<p>If the process crashes after step 1, the operation is still in the WAL. When the database starts again, it can replay the WAL and rebuild the Memtable.</p>
<pre><code class="language-text">WAL
 │
 │ replay
 ▼
Memtable
</code></pre>
<p>In my go implementation , we follow the same order.</p>
<pre><code class="language-go">wal.AddOperation("PUT", key, &amp;value)
memtable.Put(key, value)
</code></pre>
<p>Now that the write is safely recorded, we need somewhere to <strong>accumulate these writes in memory</strong>.</p>
<p>That's the job of the <strong>Memtable</strong>.</p>
<h4>Memtable</h4>
<p>Now that the write is safely recorded in the WAL, we need somewhere to keep it while it's still in memory.</p>
<p>That's the job of the Memtable.</p>
<p>A Memtable is an in-memory data structure that holds the latest writes before they are written to disk.</p>
<p>For example:</p>
<p>Memtable</p>
<p>manga:10 → Berserk
manga:11 → One Piece
manga:12 → Naruto</p>
<p>Because the data is in memory, writes are fast. We don't need to perform a disk write for every operation. The Memtable also keeps its keys sorted:</p>
<p>10 → Berserk
11 → One Piece
12 → Naruto
13 → Monster</p>
<p>This becomes useful when we eventually write the Memtable to disk.</p>
<p>What happens when it fills up?
Memory is limited, so the Memtable can't keep growing forever.
Once it reaches its size limit, we take its contents and flush them to disk as a new sorted file.</p>
<p>Memtable
    ↓
  Flush
    ↓
SSTable</p>
<p>The important part is that we don't update an existing file.
We create a new SSTable containing the Memtable's data.
After the flush, the Memtable can be cleared and start accepting new writes.</p>
<p>Memtable
   ↓
SSTable 1</p>
<p>Memtable
   ↓
SSTable 2</p>
<p>Memtable
   ↓
SSTable 3</p>
<p>This is where the LSM approach starts to become interesting: writes keep creating new immutable files instead of modifying existing files.</p>
<p>So now we have another problem:</p>
<p>What exactly is an SSTable, and how do we store and read these files?</p>
<h4>3. SSTable</h4>
<p>When the Memtable fills up, we need to move its data to disk.</p>
<p>Instead of modifying an existing file, we write the entire sorted Memtable into a <strong>new file</strong>.</p>
<p>This file is called an <strong>SSTable</strong>, or <strong>Sorted String Table</strong>.</p>
<p>For example, our Memtable might contain:</p>
<pre><code class="language-text">Memtable

10 → Berserk
11 → One Piece
12 → Naruto
13 → Monster
</code></pre>
<p>When it is flushed:</p>
<pre><code class="language-text">Memtable
    ↓
  Flush
    ↓
SSTable

10 → Berserk
11 → One Piece
12 → Naruto
13 → Monster
</code></pre>
<p>The data is already sorted, so we can write it to disk in sorted order.</p>
<p> SSTables are immutable</p>
<p>Once an SSTable is written, <strong>we never modify it</strong>.</p>
<p>Suppose we later update:</p>
<pre><code class="language-text">10 → Berserk
</code></pre>
<p>to:</p>
<pre><code class="language-text">10 → Berserk Deluxe Edition
</code></pre>
<p>We don't open the old SSTable and change it.</p>
<p>The new value goes through the same process:</p>
<pre><code class="language-text">WAL
 ↓
Memtable
 ↓
New SSTable
</code></pre>
<p>So for some time, we might have:</p>
<pre><code class="language-text">SSTable 1

10 → Berserk


SSTable 2

10 → Berserk Deluxe Edition
</code></pre>
<p>When reading, we check the newest SSTables first. The newer value therefore takes precedence.
This is one of the key differences from the B-Tree approach we saw earlier: <strong>we don't update old files in place; we keep writing new ones.</strong></p>
<p>But there is a problem</p>
<p>Every Memtable flush creates another SSTable.</p>
<p>After enough writes, we could end up with:</p>
<pre><code class="language-text">SSTable 1
SSTable 2
SSTable 3
SSTable 4
SSTable 5
...
</code></pre>
<p>Now a read may have to check multiple files.</p>
<p>Some of those files may also contain <strong>old versions of the same keys</strong>.</p>
<p>So we need a way to merge these files and remove data that is no longer needed.</p>
<p>That's where <strong>compaction</strong> comes in.</p>
<h4>4. Compaction</h4>
<p>We now have multiple SSTables on disk:</p>
<pre><code class="language-text">SSTable 1
10 → Berserk
11 → One Piece

SSTable 2
10 → Berserk Deluxe Edition
12 → Naruto

SSTable 3
13 → Monster
14 → Death Note
</code></pre>
<p>There are two problems here:</p>
<ol>
<li>A read may need to check multiple SSTables.</li>
<li>Older SSTables may contain values that have already been replaced.</li>
</ol>
<p>We can solve both by <strong>merging SSTables together</strong>.</p>
<p>This process is called <strong>compaction</strong>.</p>
<pre><code class="language-text">SSTable 1 ──┐
SSTable 2 ──┼──→ Compaction → New SSTable
SSTable 3 ──┘
</code></pre>
<p>Since each SSTable is already sorted, we can efficiently merge them.</p>
<p>For duplicate keys, the <strong>newest value wins</strong>:</p>
<pre><code class="language-text">SSTable 1
10 → Berserk

SSTable 2
10 → Berserk Deluxe Edition
</code></pre>
<p>After compaction:</p>
<pre><code class="language-text">New SSTable

10 → Berserk Deluxe Edition
</code></pre>
<p>The older version can now be discarded.</p>
<p>The same applies to deletes. Instead of immediately removing a key from older SSTables, we record a <strong>tombstone</strong> indicating that the key was deleted. During compaction, that tombstone can be used to remove the older value.</p>
<p>So instead of:</p>
<pre><code class="language-text">SSTable 1
SSTable 2
SSTable 3
SSTable 4
SSTable 5
</code></pre>
<p>we periodically reduce them to fewer, larger files.</p>
<pre><code class="language-text">SSTable 1 ──┐
SSTable 2 ──┤
SSTable 3 ──┼──→ Compaction
SSTable 4 ──┤
SSTable 5 ──┘
                 ↓
             SSTable 6
</code></pre>
<p>The goal is simple:</p>
<blockquote>
<p><strong>Keep the number of SSTables manageable and remove old versions of data.</strong></p>
</blockquote>
<p>At this point, we have the core LSM Tree write path:</p>
<pre><code class="language-text">Write
  ↓
 WAL
  ↓
Memtable
  ↓
SSTable
  ↓
Compaction
</code></pre>
<p>The next question is now about the <strong>read path</strong>: if the same key can exist in the Memtable and multiple SSTables, how do we find the correct value efficiently?</p>
<h3>Reading and Deleting Data</h3>
<p>Now that we have all the components, let's look at what actually happens when we <strong>read or delete</strong> a key.</p>
<h4>Read</h4>
<p>A read starts with the newest data and moves towards older SSTables:</p>
<pre><code class="language-text">GET manga:10
      ↓
  Memtable
      ↓
Newest SSTable
      ↓
Older SSTable
      ↓
    ...
</code></pre>
<p>If the key is found, we return the value immediately.</p>
<p>But checking every SSTable can become expensive as the number of files grows.</p>
<p>So we use a <strong>Bloom Filter</strong> for each SSTable.</p>
<p>Before reading a file, the Bloom Filter tells us whether the key <strong>might</strong> exist in that file:</p>
<pre><code class="language-text">GET manga:10
      ↓
  Memtable
      ↓
 Bloom Filter
   /       \
  No       Maybe
  ↓          ↓
Skip      Read SSTable
</code></pre>
<p>A <code>No</code> means the key definitely isn't there, so we skip the disk read.</p>
<p>A <code>Maybe</code> means we have to check the SSTable. Bloom Filters can have false positives, but they <strong>never have false negatives</strong>.</p>
<p>This makes reads much cheaper when we have many SSTables.</p>
<h4>Delete</h4>
<p>Because SSTables are immutable, we can't simply remove a key from an old SSTable.</p>
<p>Instead, a delete is written as a <strong>tombstone</strong>:</p>
<pre><code class="language-text">DELETE manga:10
</code></pre>
<p>The tombstone is stored like any other write and eventually ends up in an SSTable.</p>
<p>During a read, if we encounter the tombstone:</p>
<pre><code class="language-text">GET manga:10
      ↓
SSTable
      ↓
10 → DELETE
      ↓
  Not Found
</code></pre>
<p>We stop searching. This prevents an older SSTable from returning a value that has already been deleted.</p>
<p>During compaction, the old value and its tombstone can eventually be removed together.</p>
<hr />
<h2>Closing Thoughts</h2>
<p>Reading about LSM Trees is one thing, but implementing one makes the trade-offs much easier to understand.</p>
<p>I built a small LSM-based key-value store in Go from scratch to get a better feel for how the pieces actually fit together — from the WAL and Memtable to SSTables, compaction, tombstones, and Bloom Filters.</p>
<p>The implementation is intentionally simple and is meant for learning rather than production use.</p>
<p>You can find the complete project here:</p>
<p><a href="https://github.com/Prathamesh017/lsm-kv-store-golang">LSM KV Store in Go</a></p>
<p>The biggest takeaway for me was that an LSM Tree isn't really about one complicated data structure. It's about making a series of trade-offs:</p>
<p>Make writes cheap by avoiding in-place updates, and pay the cost later through compaction.</p>
<p>Once you look at it this way, the different pieces of an LSM Tree start to make a lot more sense.</p>
]]></content:encoded></item></channel></rss>