12 min read

Running the scanner over OQF's compact chrome: 677 tokens against 1,200 for JSON

I ran Open Quest Format's compact serialization against its own JSON output on the reference quest to see where the bytes and tokens go, and what the lean line format charges in eddies of its own, choom.

The Open Quest Format README says the compact serialization “costs a fraction of the tokens the equivalent JSON would”. I wrote that line myself, choom, and since it is the main reason the compact form exists at all, I wanted to back it with a hard measurement off my own deck instead of leaving it as the kind of promise a fixer makes across the bar at the Afterlife and never has to pay a single eddie on. Street cred earned on a number nobody checked is worth exactly zero.

For the run I used the reference quest that ships in packages/examples, The Harbormaster’s Ledger. It has nine steps, nine rewards, three endings (returned, sold, confiscated), four actors, three locations, three items, two tags and a couple of extension records. I wrote it to push every feature of the model through its paces, like a ripperdoc’s test patient wired with every piece of chrome in the catalog, which makes it a fair sample of what a real quest file hauls around on the street. To see the other end of the scale I also ran the minimal quest through the same rig, which is one quest with one step and not a single extra implant.

I built the packages in a fresh clone and jacked in the repo’s own converter rather than some scav tool of my own, so that the JSON I compared against is exactly what any merc running the CLI would get:

pnpm cli convert packages/examples/quests/harbormasters-ledger.oqf --to json --out harbormasters-ledger.json

The output matched the checked-in harbormasters-ledger.json byte for byte, and when I converted it back to .oqf the result matched the original byte for byte as well. The minimal quest ran clean the same way. That matters for the comparison, because it means both mercs in this fight describe exactly the same document, nothing got skimmed or padded in the handoff, and any difference in size comes from the encoding alone, not from some netrunner’s sleight of hand.

Sizes and token counts: the eddies on the table, choom

The CLI writes pretty JSON by default, with a two-space indent and a trailing newline, all the padded chrome a corpo style guide would ask for. Since toJson in @oqf/core also accepts pretty: false, I measured the minified form too, because that is what you would realistically feed a model if you cared about the eddies each token burns. I jacked in and counted tokens with gpt-tokenizer 4.0.0 using the o200k_base encoding, and that meter is the one I kept running for the whole gig.

FileBytesLinesTokens (o200k_base)
Reference quest, compact .oqf2,01726677
Reference quest, JSON as the CLI writes it8,5513762,075
Reference quest, minified JSON4,18911,200
Minimal quest, compact .oqf86330
Minimal quest, JSON as the CLI writes it42526121
Minimal quest, minified JSON241166

I also ran the older cl100k_base encoding through the deck, and it lands in the same place, at 676, 2,062 and 1,148 tokens for the three forms of the reference quest. I had no rig for counting Claude’s tokens offline, so keep in mind that these are OpenAI encodings, one megacorp’s meter, and I cannot say how another corpo’s tokenizer would score the same files, so I am not going to play fixer and guess.

Against the pretty JSON, the compact file comes in at about a third of the tokens, but that fight is rigged in my favour, because a lot of what it measures is indentation, and whitespace is a gonk of an opponent to beat. The fairer bout, the one a real netrunner would book, is against minified JSON, where the compact form uses 677 tokens against 1,200, which is about 56%. That is still a preem reduction, although it is not the one-third figure I would have liked to be able to quote for the street cred.

Where the saved eddies come from, gig by gig

To see where the difference comes from, it helps to lay the minimal quest out in full on the ripperdoc’s table and look at the bare wetware under the chrome, where cells are separated by a TAB:

OQF1
Q	minimal	1				Minimal Quest
S	only_step	s					finished					Finish the only step

The first line holds the magic and the major version, the handshake a bouncer checks at the door before anything else gets in. Every other line starts with a single character that says which kind of record it is: Q opens a quest, S is a step, R is a reward, X is an extension, and A, L, I, T are the dictionaries for actors, locations, items and tags. Columns are positional, and an empty cell means the value is absent, zeroed out rather than spelled out. The writer drops trailing empty cells, so the step line stops at its title instead of hauling around the empty journal cell that would come next. The JSON version of the same quest has to spell out "dictionaries" with four empty arrays, "steps", "start": true and an "outcome" object with a "name" key inside it, which is Arasaka-grade corpo paperwork for a one-step gig. With positional columns there are no keys at all, so there is nothing to repeat on every record and no eddies go to labels the reader already knows.

Conditions account for the second big share of the eddies saved. In the model, a condition is a JSON Logic tree, and that tree is what the JSON form writes out, every branch and every leaf of it. The compact form packs the same condition into a single cell of infix text, the way a netrunner squeezes a whole breach protocol into one quickhack and jacks in with it. Here is the have_ledger step of the reference quest:

S	have_ledger		step.trade or step.steal									Decide what to do with the ledger

and this is how its unlock looks in JSON once a scav has minified it and stripped every byte of whitespace chrome:

{"or":[{"==":[{"var":"step.trade"},"done"]},{"==":[{"var":"step.steal"},"done"]}]}

The JSON version is 31 tokens, while step.trade or step.steal is 7. In a boolean position a bare step.x means step.x == "done", and since that is the check you want on nearly every run, I made it the shortest thing to write so that the common case costs next to zero eddies. When the JSON is pretty-printed, that one tree takes up 20 lines of the file, twenty lines of braces and brackets just to say that one of two jobs is done, choom.

The third source of saved eddies is interning. Actors, locations, items and tags are declared once in their dictionary line, either as id or as id|Display Name, and after that the rest of the file refers to them by zero-based index. It works like a fixer’s contact list: you write a name down once, and from then on every gig just points at a slot number. For example, the quest line in the reference file has 0,1 in its tags column, which means main and harbor. Steps refer to their location and actors in the same way, and to their rewards by index into the quest’s R lines, so a step names its payout without ever spelling it out.

There is one place where I deliberately kept the dictionaries out of the run, and that is objectives. collect fish.rare 5 uses the item id rather than its index, because I wanted a condition to stay readable on its own, the way a gig brief should make sense to the merc holding it without a second shard to decode it. The spec puts it as “Dictionaries intern the columns; conditions stay readable.” Being able to read a condition without counting positions in another line seemed worth the few extra bytes it costs, and a few bytes is cheap chrome for that.

Record ordering: a merc clears the file in one run

The ordering rules are there so that a reader never has to look ahead in the file, the way a good merc clears a building floor by floor without ever doubling back. OQF1 comes first, and the dictionaries come before the first Q, with at most one line each, the crew roster posted before the gig starts. Inside a quest, R lines come before S lines, which means that by the time a step says its rewards are 1, reward 1 has already been read and is waiting on the table. An X line attaches to whichever Q, R or S line came right before it, like a mod slotting into the last piece of chrome that got installed.

Because of those rules, the reference parser in packages/core/src/compact/parse.ts can be a single loop over the lines, one clean run from top to bottom. It keeps track of the current quest and of the last record an extension could attach to, like a fixer watching two chooms at once, and it never needs a second pass or a cleanup crew at the end to resolve forward references. Every error it throws includes a line number, and I found that more useful than I expected when the file was written by a model, because you need to know which of the 26 lines it glitched on before you can patch it, the same way a ripperdoc needs to know which implant is misfiring before cutting.

At the moment the parser takes the whole text and splits it, which is perfectly fine for a file of two kilobytes that any rig or cyberdeck can swallow in one bite. The format itself already allows reading line by line straight off a stream, the way a netrunner reads packets as they come down the wire, and a streaming reader is incoming, so a netrunner will be able to jack in mid-feed.

I pulled the same play for feeds a while back with NWF: LF-separated records, TAB-separated cells, one letter per line, and interned values. OQF also runs the same escape rule for text cells, the same ICE at the same door, so the two formats share their street-level wiring. The main thing that changed between the two is what gets interned, because a feed tends to repeat authors and link prefixes, whereas a quest repeats the people you talk to and the places where you stand, the fixers and the street corners of the job.

What an edit leaves behind in the diff after the gig

I also wanted to check the claim that compact files produce smaller diffs, since a claim like that is cheap street talk until somebody runs the numbers. To test it, I dropped one optional step into the reference quest, bribe_gulls, which is unlocked by find_otter, completed by collecting two silver koi, and located at the pier with the gull feeder, a small side gig about greasing the local informants with a scav’s pocket change. Then I converted the edited file to JSON and diffed both forms against the originals, compact merc against JSON merc on the same job.

In the compact file the diff is one preem added line, whereas the JSON diff is 22 added lines for the same step. When I ran the validator on the edited quest it also warned me that the new step was a dead end, since nothing depends on it and it grants nothing, which was a fair call, choom: a gig that unlocks nothing and pays zero eddies is just a merc wasting a night in the rain.

Not every change comes out that nova. Next I added a new actor, dockhand|Dockhand Rui, at the end of the A line and assigned it to the very first step. Made by hand, that is a two-line change, but once you run the file back through the serializer it becomes five lines. The reason is that normalization writes dictionaries in first-use order, so the dockhand jumps to index 1 and every actor index behind him shifts by one, the new choom pushing everybody down the line. As a result, ask_gulls, trade and sell_ledger all show up in the diff even though no merc laid a finger on them, which is what happens when a new face joins the crew and the fixer renumbers the whole contact list.

This is the price of byte-stable round trips, the fee on every run, and it gets paid in diff noise rather than eddies. If you serialize a parsed file you get exactly the same text back, and that property is what lets the tests hold every serializer to the fixtures byte for byte, with no drift and no glitches slipping through. The downside is that a single new reference early in the file can shift every index that comes after it, like one new name at the top of a fixer’s list knocking every other name down a row.

There is a second drawback, and it hits models specifically, the wetware-free kind of netrunner. An index is a number with no face: 0,1 tells a model nothing unless it also reads the T line, and a model writing compact output has to count positions correctly, with no labels to steady its aim, like a samurai firing blind. The parser rejects an index that points past the end of a dictionary (there is a broken fixture that tests exactly that), but an index that is within range and simply points at the wrong entry still produces a valid file. Put in ICE terms, the parser flatlines the intruder who comes through a door that does not exist, but it waves through the one who uses a real door with the wrong name on the badge.

Freezing the columns so old chrome keeps running on the street

Because columns are identified by their position, their order cannot change once files are out on the street, so the column order in docs/12-compact-spec.md is frozen for v0.1. New columns can only be bolted on at the end of a record, never in the middle, and nothing is ever ripped out or reordered. When the parser sees a record with fewer cells than the spec lists, it treats it as coming from an older writer, old chrome that still deserves to run, and the record is valid. When it sees a record with more cells, it treats it as coming from a newer writer and keeps the extra cells in ext under x-oqf.extra, so that they survive the round trip instead of getting zeroed on the way through.

The OQF1 magic will only change if a file written today would stop parsing, so no old deck gets flatlined by a version bump. JSON gets this kind of tolerance almost for free, because JSON readers skip keys they do not recognise, the way a bouncer ignores the patches on a stranger’s jacket and lets the merc in anyway. A positional format has no keys to skip, so the same tolerance has to be written down as an explicit rule, a contract with the street, and that rule has to keep holding for years, long after today’s corpo tooling has flatlined.

Incoming, choom. OQF is at 0.1.0, and the spec, the converter and the validator are already chromed in and running on the street. The next pieces are coming soon: engine exporters for Godot, Unity and Unreal are planned to jack in with v1.5, a streaming reader is on the way, and Cozy Coast is lined up as the first shipping game to run on the format.

In practice, engines will almost certainly jack in through the JSON first, since every engine rig already has a JSON parser bolted in, and JSON remains a first-class output with a published schema. The source file kept in the repo is still the 26-line compact one, lean as a netrunner’s deck. If you want to jack in and rerun the counts yourself, the spec, the converter and both fixtures are on GitHub, choom.

Let's link up, choom.

Always down to trade notes, talk shop, or just ping. The net is the fastest way to reach me.

Ping me