9 min read

Measuring the OQF compact format: 677 tokens against 1,200 for JSON

I measured Open Quest Format's compact serialization against its own JSON output on the reference quest to see where the bytes and tokens go, and what the line format costs in return.

The Open Quest Format README says the compact serialization “costs a fraction of the tokens the equivalent JSON would”. I wrote that sentence myself, and since it is the main reason the compact form exists at all, I wanted to back it with an actual measurement instead of leaving it as a claim in a README.

For the test I used the reference quest that ships in packages/examples, The Harbormaster’s Ledger. It has nine steps, nine rewards, three endings (returned, sold, confiscated), four actors, three locations, three items, two tags and a couple of extension records. I wrote it to exercise every feature of the model, which makes it a fair sample of what a real quest file would contain. To see the other end of the scale I also ran the minimal quest, which is one quest with one step and nothing else.

I built the packages in a fresh clone and ran the repo’s own converter, so that the JSON I compared against is exactly what someone using the CLI would get:

pnpm cli convert packages/examples/quests/harbormasters-ledger.oqf --to json --out harbormasters-ledger.json

The output matched the checked-in harbormasters-ledger.json byte for byte, and when I converted it back to .oqf the result matched the original byte for byte as well. The minimal quest behaved the same way. That matters for the comparison, because it means both sides describe exactly the same document and any difference in size comes from the encoding alone.

Sizes and token counts

The CLI writes pretty JSON by default, with a two-space indent and a trailing newline. Since toJson in @oqf/core also accepts pretty: false, I measured the minified form too, because that is what you would realistically send to a model if you cared about tokens. I counted tokens with gpt-tokenizer 4.0.0 using the o200k_base encoding.

FileBytesLinesTokens (o200k_base)
Reference quest, compact .oqf2,01726677
Reference quest, JSON as the CLI writes it8,5513762,075
Reference quest, minified JSON4,18911,200
Minimal quest, compact .oqf86330
Minimal quest, JSON as the CLI writes it42526121
Minimal quest, minified JSON241166

I also tried the older cl100k_base encoding, and it lands in the same place, at 676, 2,062 and 1,148 tokens for the three forms of the reference quest. I did not have a way to count Claude’s tokens offline, so keep in mind that these are OpenAI encodings and I cannot say how other tokenizers would score the same files.

Against the pretty JSON, the compact file comes in at about a third of the tokens, but that comparison is generous to my side because a lot of what it measures is indentation. The fairer comparison is against minified JSON, where the compact form uses 677 tokens against 1,200, which is about 56%. That is still a large reduction, although it is not the one-third figure I would have liked to be able to quote.

Where the savings come from

To see where the difference comes from, it helps to look at the minimal quest in full, where cells are separated by a TAB:

OQF1
Q	minimal	1				Minimal Quest
S	only_step	s					finished					Finish the only step

The first line holds the magic and the major version. Every other line starts with a single character that says which kind of record it is: Q opens a quest, S is a step, R is a reward, X is an extension, and A, L, I, T are the dictionaries for actors, locations, items and tags. Columns are positional, and an empty cell means the value is absent. The writer drops trailing empty cells, so the step line stops at its title instead of writing out the empty journal cell that would come next. The JSON version of the same quest has to spell out "dictionaries" with four empty arrays, "steps", "start": true and an "outcome" object with a "name" key inside it. With positional columns there are no keys at all, so there is nothing to repeat on every record.

Conditions account for the second big share of the savings. In the model, a condition is a JSON Logic tree, and that tree is what the JSON form writes out. The compact form expresses the same condition as a single cell of infix text. Here is the have_ledger step of the reference quest:

S	have_ledger		step.trade or step.steal									Decide what to do with the ledger

and this is how its unlock looks in JSON once minified:

{"or":[{"==":[{"var":"step.trade"},"done"]},{"==":[{"var":"step.steal"},"done"]}]}

The JSON version is 31 tokens, while step.trade or step.steal is 7. In a boolean position a bare step.x means step.x == "done", and since that is the check you want nearly every time, I made it the shortest thing to write so that the common case costs almost nothing. When the JSON is pretty-printed, that one tree takes up 20 lines of the file.

The third source of savings is interning. Actors, locations, items and tags are declared once in their dictionary line, either as id or as id|Display Name, and after that the rest of the file refers to them by zero-based index. For example, the quest line in the reference file has 0,1 in its tags column, which means main and harbor. Steps refer to their location and actors in the same way, and to their rewards by index into the quest’s R lines.

There is one place where I deliberately kept the dictionaries out, and that is objectives. collect fish.rare 5 uses the item id rather than its index, because I wanted a condition to stay readable on its own. The spec puts it as “Dictionaries intern the columns; conditions stay readable.” Being able to read a condition without counting positions in another line seemed worth the few extra bytes it costs.

Record ordering and the single-pass parser

The ordering rules are there so that a reader never has to look ahead in the file. OQF1 comes first, and the dictionaries come before the first Q, with at most one line each. Inside a quest, R lines come before S lines, which means that by the time a step says its rewards are 1, reward 1 has already been read. An X line attaches to whichever Q, R or S line came right before it.

Because of those rules, the reference parser in packages/core/src/compact/parse.ts can be a single loop over the lines. It keeps track of the current quest and of the last record an extension could attach to, and it never needs a second pass or a cleanup step at the end to resolve forward references. Every error it throws includes a line number, and I found that more useful than I expected when the file was written by a model, because you need to know which of the 26 lines it got wrong.

At the moment the parser takes the whole text and splits it, which is perfectly fine for a file of two kilobytes. The format itself already allows reading line by line from a stream, and a streaming reader is incoming.

I took the same approach for feeds a while back with NWF: LF-separated records, TAB-separated cells, one letter per line, and interned values. OQF also uses the same escape rule for text cells. The main thing that changed between the two is what gets interned, because a feed tends to repeat authors and link prefixes, whereas a quest repeats the people you talk to and the places where you stand.

How edits show up in a diff

I also wanted to check the claim that compact files produce smaller diffs. To test it, I added one optional step to the reference quest, bribe_gulls, which is unlocked by find_otter, completed by collecting two silver koi, and located at the pier with the gull feeder. Then I converted the edited file to JSON and diffed both forms against the originals.

In the compact file the diff is one added line, whereas the JSON diff is 22 added lines for the same step. When I ran the validator on the edited quest it also warned me that the new step was a dead end, since nothing depends on it and it grants nothing, which was a fair point.

Not every change produces such a small diff. Next I added a new actor, dockhand|Dockhand Rui, at the end of the A line and assigned it to the very first step. Made by hand, that is a two-line change, but once you run the file back through the serializer it becomes five lines. The reason is that normalization writes dictionaries in first-use order, so the dockhand moves to index 1 and every actor index after it shifts by one. As a result, ask_gulls, trade and sell_ledger all show up in the diff even though nobody edited them.

This is the price of byte-stable round trips. If you serialize a parsed file you get exactly the same text back, and that property is what lets the tests hold every serializer to the fixtures byte for byte. The downside is that a single new reference early in the file can shift every index that comes after it.

There is a second drawback, and it affects models specifically. An index does not mean anything by itself: 0,1 tells a model nothing unless it also reads the T line, and a model writing compact output has to count positions correctly. The parser rejects an index that points past the end of a dictionary (there is a broken fixture that tests exactly that), but an index that is within range and simply points at the wrong entry still produces a valid file.

Keeping the column layout compatible

Because columns are identified by their position, their order cannot change once files exist in the wild, so the column order in docs/12-compact-spec.md is frozen for v0.1. New columns can only be added at the end of a record, never in the middle, and nothing is ever removed or reordered. When the parser sees a record with fewer cells than the spec lists, it treats it as coming from an older writer, and the record is valid. When it sees a record with more cells, it treats it as coming from a newer writer and keeps the extra cells in ext under x-oqf.extra, so that they survive the round trip.

The OQF1 magic will only change if a file written today would stop parsing. JSON gets this kind of tolerance almost for free, because JSON readers skip keys they do not recognise. A positional format has no keys to skip, so the same tolerance has to be written down as an explicit rule, and that rule has to keep holding for years.

Incoming. OQF is at 0.1.0, and the spec, the converter and the validator are already in place. The next pieces are coming soon: engine exporters for Godot, Unity and Unreal are planned to arrive in v1.5, a streaming reader is on the way, and Cozy Coast is lined up as the first shipping game to run on the format.

In practice, engines will almost certainly import the JSON first, since every engine already has a JSON parser, and JSON remains a first-class output with a published schema. The source file kept in the repo is still the 26-line compact one. If you want to rerun the counts yourself, the spec, the converter and both fixtures are on GitHub.

Let's connect.

Always happy to talk shop, compare notes, or just say hi. Email or LinkedIn is the fastest way to reach me.

Get in touch