Munkherdene

I Spent Three Years to Understand This Tiny Program

Munkherdene ยท August 29, 2026 ยท โ˜• 11 min read ยท Season 1 ยท Cecile

I Spent Three Years to Understand This Tiny Program

Building my own programming language โ€” Season 1: Cecile ยท Part 1: The Interpreter

Here is the first program that ever took me three years to understand:

println "Good day";
Good day

One line. One greeting. Three years.

I don't mean I stared at this file for three years. I mean it took me that long to honestly answer a question hiding inside it: when you hand a machine a few letters like these, how does it actually run them? Not the hand-wavy version โ€” the real one, all the way down.

println "Good day";?CPU
How does a machine actually run a few letters?

This is where that question started. Over one year I built my own little programming language, called Cecile, and in doing so I discovered that the most boring program you know โ€” "hello world" โ€” is quietly hiding almost everything interesting about computers.

So โ€” why do this at all?

I was a second-year student who hadn't figured anything out yet.

I kept running into the same 1am thought: how am I going to get a job, and what do I actually have that sets me apart from the next student? I had zero work experience. So I decided, in the slightly desperate way students decide things, that I needed one real thing I could point at and say: I built that.

how do I get a job?what sets me apart?01:00

So I went looking for a project.

The first goal was honestly that shallow โ€” find something impressive enough to help me get hired. I was scrolling through the build-your-own-x repo (a big list of "build your own git / database / regex engine / โ€ฆ" guides) when one line stopped me cold: build your own compiler / interpreter.

I didn't fully know what those words meant. That was the point. It just looked interesting.

"Just because it looked interesting?"

Yes. And because it smelled like the kind of project where you accidentally learn ten other things on the way to the one thing.

Here's the belief that pulled me in, and I still hold it: the way code runs is the most mysterious, most fundamental piece of software there is. It grew up right alongside the computer itself. Strip it away and the word "software" stops meaning anything โ€” there's nothing left to write software in. Every app, every website, every game you've ever touched sits on top of some program whose only job is to take human-ish text and turn it into something a chip will do. I had used that machinery every day for years and had no idea how it worked.

At some point, I figured, you have to look inside.

I also knew, without being able to explain why yet, that this was going to hurt. Compilers sit on top of half of computer science at once โ€” grammars, parsing, data structures, memory, a little of how the CPU itself thinks. Which was exactly the appeal. Hard thing, huge surface area, lots to learn.

For something that hard, I didn't want to flail โ€” I wanted structure. So I did the un-glamorous thing and bought a book.

The book: Crafting Interpreters

Crafting Interpreters by Robert Nystrom (finished in 2020) deserves every bit of praise it gets. It's beginner-friendly without being shallow, it's full of hand-drawn diagrams, and โ€” rare for a technical book โ€” it's genuinely fun to read. If you want to learn this topic, start here. I mean it.

The format is simple: you don't read about interpreters, you build one, chapter by chapter. Each chapter hands you code, an explanation, and pictures, and by the end you have a real, working language. The book comes in two halves, and the jump between them is the whole education.

Part 1 โ€” A Tree-Walk Interpreter (10 chapters). You build a small language called Lox in Java. It's called a tree-walk interpreter because it literally builds a tree out of your program and then walks that tree, node by node, doing what each node says. It's the most direct way to make a language work. Most beginners end up here โ€” and that's completely fine. I did.

Part 2 โ€” A Bytecode Virtual Machine (17 chapters). This is the heart of the book. You throw away the Java version and rebuild the language from scratch in C, this time compiling it to bytecode and running that on a little virtual machine you write yourself. It's a huge step up, and it's where Cecile really took shape โ€” enough that it gets its own post next. For now, just know that this is the model real languages actually use: JavaScript's V8, Python, and Java all run on some version of this bytecode-VM idea.

By the end of the book, I didn't just have a working tutorial language. I had opinions.

Designing a language of my own: Cecile

So I decided I didn't want to build Lox. I wanted mine. I called it Cecile.

At the time I was learning Rust, and I'd fallen for it hard โ€” a genuinely wonderful language (the borrow checker and I have our differences, but that's a story for another day). So I asked myself one deceptively small question:

What if it read like Rust, behaved like JavaScript, and had types like TypeScript?

That became Cecile's whole personality. Rust's clean, explicit syntax โ€” fn, let, arrows for return types. JavaScript's easygoing runtime feel. TypeScript's "types are there when you want them." I wrote it in Rust, taking early inspiration from loxcraft, a Rust implementation of Lox. A program in Cecile looks like this:

fn add(a: int, b: int) -> int {
    return a + b;
}

let total: int = add(2, 3);
println total;        // 5

But taste doesn't run on a computer. To actually build a language, you start by writing down its grammar โ€” the exact rules for what counts as a valid program. Grammar is the contract. Once you have it, everything downstream is a series of translations, each one turning your program into a form that's a little closer to something a machine can execute. Let me show you those stages by pushing one line of Cecile through all of them.

Watch one line become something a machine can run

Say a Cecile program contains this:

let x: int = 2 + 3 * 4;

Watch where it goes.

Stage 1 โ€” the Lexer: text โ†’ tokens

A computer doesn't see words. It sees a river of characters: l, e, t, space, x, :, and so on. The lexer (or scanner) is the first translator. Its only job is to chop that river into meaningful chunks called tokens, and label each one.

So let x: int = 2 + 3 * 4; comes out the other side as:

[let โ†’ keyword]  [x โ†’ identifier]  [: โ†’ colon]  [int โ†’ type]  [= โ†’ equals]
[2 โ†’ number]  [+ โ†’ plus]  [3 โ†’ number]  [* โ†’ star]  [4 โ†’ number]  [; โ†’ semicolon]

That's it. Text in, labeled tokens out. The lexer doesn't understand math or meaning yet โ€” it doesn't know that * binds tighter than +. It just knows "this blob is a number, that blob is a keyword, this symbol is a star." Dumb, fast, and absolutely necessary.

Stage 2 โ€” the Parser: tokens โ†’ tree

This is the hard part. Honestly, one of the hardest steps in the whole project.

The parser takes that flat line of tokens and discovers the structure hiding in it โ€” the shape my grammar says it must have. There's a whole zoo of parsing algorithms, and I went with Pratt parsing on top of recursive descent. The reason I reached for Pratt is exactly the thing the lexer couldn't handle: precedence. 2 + 3 * 4 is not (2 + 3) * 4 = 20. It's 2 + (3 * 4) = 14, because multiplication binds tighter than addition โ€” and the parser is where that rule finally gets enforced.

The output isn't a line anymore. It's a tree:

=x+2*34evaluated first:3 * 4 = 12then 2 + 12 = 14
The syntax tree for x = 2 + 3 * 4

Read it and the shape carries the meaning: to compute +, you first have to compute *, because * lives below it in the tree. The tree's shape solves the ordering problem. This tree has a proper name โ€” an Abstract Syntax Tree, or AST โ€” and it's the most satisfying artifact in the whole pipeline, because it's the moment raw text turns into something that actually knows what it means. (If the tokens don't fit the grammar โ€” say you wrote let x = = 2 โ€” this is exactly where the parser stops and hands you a syntax error.)

Stage 3 โ€” Bytecode generation: tree โ†’ instructions

A tree is lovely to reason about but awkward to execute fast. So the last translator walks the AST and flattens it into bytecode: a straight, ordered list of tiny instructions the machine can run one after another.

Conceptually, our little tree compiles down to something like:

PUSH 3
PUSH 4
MUL          ; 3 * 4  -> 12
PUSH 2
ADD          ; 2 + 12 -> 14
STORE x      ; x = 14

Notice how the tree's nesting became a sequence: multiply first, then add, then store. That's the tree's meaning, unrolled into steps a machine can follow one at a time without ever chasing a pointer through a tree again.

Text โ†’ tokens โ†’ tree โ†’ instructions. Four translations, and Cecile has a runnable program. (What actually runs that bytecode is a little machine of its own โ€” that's the next post.)

source textlet x: int = 2 + 3 * 4;Lexertokens[let] [x] [:] [int] [=] [2] [+] โ€ฆParsertree (AST)x = 2 + (3 * 4)Bytecode generatorbytecodePUSH 3 PUSH 4 MUL PUSH 2 ADD STORE xVirtual machine(next post)resultx = 14
The pipeline: lexer, parser, bytecode generator, then the VM (next post)

Wait โ€” interpreter or compiler? What's the difference?

Good moment to stop, because you might be quietly wondering what I was wondering the whole first month: what even is an interpreter, and how is it different from a compiler?

Here's the analogy that finally made it click for me, and I've never found a better one:

An interpreter is a live translator. A compiler is a document translator.

Interpretera live translatorCompilera document translatorline 1runline 2runline 3runstarts now, one line at a timeline 1line 2line 3read thewhole textfinishedtranslationwaits, sees everything, optimizes
Interpreter vs compiler

A live translator works line by line, sentence by sentence, translating on the spot as you speak. Immediate, flexible, no waiting โ€” but they only ever see one sentence at a time. A document translator does the opposite: they take the whole text first, read all of it, analyze it, understand how the end connects to the beginning, and only then produce the translation.

Put that way, the difference is huge โ€” and so is the trade-off. The document translator (the compiler) can produce faster, more optimized output, precisely because it sees everything before committing. The live translator (the interpreter) gives that up in exchange for starting right now, with no wait and more flexibility. Neither is "correct"; they're different deals. Cecile is an interpreter โ€” it compiles your code to bytecode and then immediately runs it, no separate build step.

That's the key point for today: even a plain interpreter gives us exactly the result we want, and it's more than enough to finally see how running code works. Which is why Cecile's whole first year lived here, in software, one comfortable step above the metal.

An aside: so why isn't there just one language that does everything?

This question nagged at me for months, so let me hand you the mental model that finally put it to rest.

Learning to drive a car is one thing. Understanding the engine โ€” the parts, how they connect, why the thing moves โ€” is a completely different thing. And depending on what engine and parts a car is built from, you get a completely different vehicle: a little economy hatchback, a cargo truck, a piece of heavy machinery. Programming languages are the same. Their design, structure, and purpose pull them in different directions, and no single design is best at everything. (It goes deeper than taste, too โ€” all the way down to the chip. But that's a rabbit hole for a later post.)

If you find that you're spending almost all your time on theory, start turning some attention to practical things; it will improve your theories. If you find that you're spending almost all your time on practice, start turning some attention to theoretical things; it will improve your practice. โ€” Donald Knuth

Does it actually run? Here's Cecile, working

Enough theory. After all of it โ€” grammar, lexer, parser, bytecode โ€” here's a real, complete Cecile program. A number-guessing game:

fn guessGame() {
    println "Enter your guess:";

    let target: int = randomInt(100);
    let tries: int = 0;
    let maxTries: int = 10;

    while tries < maxTries {
        let guess: int = readInt();
        tries = tries + 1;

        if guess == target {
            println "congrats, you got it ๐ŸŽ‰";
            println "Your tries:";
            println tries;
            return;
        }

        if guess < target { println "too low"; }
        if guess > target { println "too high"; }
    }

    println "Out of tries! It was:";
    println target;
}

// Cecile picks a random number; the player tries to guess it.
guessGame();

Look at everything in there: typed variables, a while loop, if conditions, function calls, a random-number call, reading input from the keyboard, an emoji in a string. Every one of those features had to be lexed into tokens, parsed into a tree, lowered to bytecode, and executed โ€” by an interpreter I wrote. The first time this ran and printed too low at me, I just sat there grinning. It still feels a little unreal to paste it out.

Enter your guess:50too low

But something was still missing

And yet.

The moment Cecile worked, I realized I still couldn't answer the question I started with. Yes โ€” a handful of letters go in, and my machine produces exactly the output I asked for. But my interpreter is itself a program written in another language, running on top of another system, and somewhere underneath all of that a physical CPU is flipping actual electrical signals. How do a few letters become that? That floor was still a black box, and I'd only built one story above it.

Cecile runs โ€” but it never once touches the metal. And I couldn't stop thinking about the metal.

Next in Season 1

Before we go down, we go in. Cecile compiles your code to bytecode โ€” but what actually runs that bytecode? It turns out to be a tiny machine of its own: a loop, a stack, and a surprising amount of bookkeeping to clean up after you.

Next post: we build the virtual machine that runs Cecile. Bring a stack. ๐Ÿ‘‹