flanger

A small language frontend in Rust. Built to learn compilers by keeping the scope small.

Why, given fshell already parses

fshell already has a lexer, a parser, and a syntax tree, so starting over looks redundant. It is, on purpose. fshell's frontend grew around shell constraints: bare words, pipelines, POSIX fallback, aliases that expand as you type. Every decision there serves the shell.

flanger is the opposite exercise: a conventional language built from nothing, small enough that every part can be understood fully and tested exactly, with nothing to serve except the language itself. It is slower than reusing fshell's frontend because learning is the point, not the fastest route to a working parser. And it heads somewhere fshell never needed to go: past parsing into code generation.

The language today

This is close to the whole language right now:

fn main() -> i64 {
  return 1 + 2 * 3;
}

Functions take no parameters and hold a single return statement. The return type can be left out, in which case it defaults to Unit. Expressions are integer arithmetic with the usual precedence, and parentheses override it. Types are i64 and i32.

Lexer

Handwritten and byte based, ASCII only. UTF-8 handling is still an open todo in the code. Every token carries a span with start and end offsets, so errors point at the exact characters involved.

It tells minus apart from arrow by peeking one byte ahead, recognizes fn and return while leaving longer names like returnValue as identifiers, and reports two error kinds: unexpected character and integer overflow. Both carry spans.

Parser and errors

Recursive descent over the token stream. Additive expressions sit over multiplicative over primaries, which is what gives 1 + 2 * 3 the standard grouping, with parenthesized expressions as primaries.

Every error carries a span plus what was expected against what actually arrived, and unknown type names are their own error rather than a generic mismatch. Lex errors pass through the parser wrapped, so one error type covers both stages.

The expression tree starts with two cases:

pub enum Expression {
    Integer(i64),
    Binary(Box<BinaryExpression>),
}

Tokens, spans, and error types live in their own modules so each stage can be tested on its own. Function names borrow from the source text instead of copying it.

Testing

Fifteen tests, all passing, roughly one per behavior: token spans, minus versus arrow, keyword prefixes, overflow, precedence, associativity, parentheses. The tests pin spans exactly, not just token kinds, since spans are the diagnostic contract the rest will rely on.

Status

Frontend work in progress. Code generation and execution are not started, and the command line entry point is still a stub that prints a placeholder. The natural next steps are function parameters, more statement forms, and then a decision on a simple backend.