Skip to main content

Open format · Draft

ISOLDA

Indexed Subject-Ordered Linked Data for AI

A way of writing knowledge graphs so that language models — even small ones running on a laptop — can read them. Fewer tokens than Turtle or JSON-LD, readable names instead of opaque codes, and not a single fact lost on the way.

Work in progress, reported honestly. Version 0.1 met two of its three goals in a pre-registered study: it is lossless and cheaper, but it was not yet as accurate as the best standard format. Version 0.2 fixes the causes and is being validated now. Good and bad results are published alike.

01 · The problem

Knowledge graphs are written for machines, not for language models

More and more of what we know is published as a knowledge graph: the JSON-LD (opens in a new window) hidden in almost every web page, open data catalogues, Wikidata. The standard way to hand one to a language model is to paste it into the prompt as text, usually in Turtle (opens in a new window) or JSON-LD.

Those formats were designed for parsers, which never get tired and never guess. A language model reads differently, and it shows:

  • They are expensive. Every repeated key, prefix and bracket is a token you pay for and a piece of the context window you lose.
  • They talk in codes. Entities are named by identifiers like ex:o7 or @p3. Ask a small model who someone works for, and it often answers the code instead of the name.
  • They hide the counts. “How many datasets does this catalogue have?” means counting hundreds of scattered statements — something models do badly.

The format matters more than it seems. Google researchers found that changing only how a graph is written changed model performance by between 4.8 % and 61.8 %, depending on the task (Fatemi et al., ICLR 2024 (opens in a new window)). On the KG-LLM-Bench (opens in a new window) benchmark, the two web standards came last:

Average accuracy by format on KG-LLM-Bench (published results)

  • Structured JSON : 0.42
  • List of edges : 0.41
  • YAML : 0.41
  • Turtle : 0.35
  • JSON-LD : 0.34
Source: Markowitz et al., KG-LLM-Bench (2025), average across models and tasks. Turtle and JSON-LD were also the most expensive prompts: 8,171 and 13,503 tokens on average, against 2,645 for the list of edges — which is cheap but cannot represent RDF without loss.

So there is a gap: the formats that keep everything RDF can express (links, data types, languages) are the ones models read worst, and the ones models read best throw part of it away.

02 · The idea

Write the graph the way a model reads a table

ISOLDA keeps everything that makes RDF what it is and changes only the layout. The name spells out the design: every entity is Indexed by a readable id, statements are Subject-Ordered (all together, one row per entity), it is Linked Data (decoding it gives back exactly the same graph), and it is written for AI.

Names, not codes

Each entity is identified by its own label — “Opendata.cat”, not ex:o7. When a question asks where Anna works, the answer is already written next to her name.

One row per thing

Everything about an entity sits on one line, and entities of the same type share a table with a single header. No repeated keys, no scattered statements.

Counts written out

Every table and every list says how many items it has (Person[2], [5] a | b | …). Models are bad at counting long lists; ISOLDA does the counting for them.

A glossary that fits

Each document starts with a few lines explaining only the notation it actually uses, so even a small model knows how to read it.

The same four facts, two ways

Two people and the organizations they belong to. A simplified illustration, not byte-exact encoder output:

Turtle
@prefix schema: <http://schema.org/> .
@prefix ex: <https://example.org/> .

ex:p1 a schema:Person ;
    schema:name "Anna Puig" ;
    schema:memberOf ex:o7 .
ex:p2 a schema:Person ;
    schema:name "Pau Serra" ;
    schema:memberOf ex:o9 .
ex:o7 a schema:Organization ;
    schema:name "Opendata.cat" .
ex:o9 a schema:Organization ;
    schema:name "Sorensen.ai" .
ISOLDA
# Entities are grouped by type.
# `Type[N]{a, b}:` is a table of N entities, one per row;
#   the first cell is the entity id.
# A property ending in `>` links to another entity by its id.

@vocab http://schema.org/
@types Person 2, Organization 2

Person[2]{name, memberOf>}:
  Anna Puig, Opendata.cat
  Pau Serra, Sorensen.ai

Organization[2]{name}:
  Opendata.cat
  Sorensen.ai

In Turtle, a model has to jump from ex:p1 to ex:o7 and then look up what ex:o7 is called. In ISOLDA, Anna Puig and Opendata.cat are on the same line.

And it is lossless: the reference implementation checks that RDF → ISOLDA → RDF gives back an identical graph, IRIs, data types, language tags and blank nodes included. The output is deterministic, byte for byte.

03 · What we have measured

Fewer tokens, better counting, names that stick

ISOLDA was tested with four open-weight models that anyone can run — two small (Qwen3 4B and Gemma 4 E4B) and two mid-sized (Qwen3.8 27B and Gemma 4 31B) — with reasoning turned off, so we measure how well each model reads the format.

Tokens needed for the same graph (lower is better)

  • TOON / JSON-LD : 12,075
  • JSON-LD : 11,127
  • Turtle : 11,058
  • ISOLDA : 8,702 · −21 %
Development graph: the 43 JSON-LD blocks of xaviviro.com, 810 triples, counted with the o200k tokenizer. ISOLDA 0.1 with its default options, glossary included. “TOON / JSON-LD” is TOON applied directly to the JSON-LD, shown because ISOLDA's first prototypes started from it.

Token saving against Turtle with labels, on a graph never seen during design

  • Qwen3 4B : −20.3 %
  • Qwen3.8 27B : −25.9 %
  • Gemma 4 E4B : −28.6 %
  • Gemma 4 31B : −28.6 %
Pre-registered confirmatory run on the open data catalogue of the Generalitat de Catalunya (DCAT, 1,292-triple sample). Tokens counted with each model's own tokenizer. The pre-registered target was at least 15 %.

Small models stop answering with codes

The clearest finding so far: small models answer with whatever identifier they see. Give them codes and they reply with codes; even readable slugs like sorensen-ai come back as the answer. Using the entity's own name as its identifier fixes most of it.

Multi-hop questions answered correctly, out of 14, by how entities are identified

Qwen3 4B

  • JSON-LD : 11/14
  • Opaque ids : 3/14
  • Readable slugs : 5/14
  • ISOLDA (names) : 12/14

Gemma 4 E4B

  • JSON-LD : 12/14
  • Opaque ids : 8/14
  • Readable slugs : 3/14
  • ISOLDA (names) : 13/14
Exploratory pilot runs on the development graph, 14 frozen multi-hop questions (“who is the author of…”, “which organization…”). Small numbers and no confidence intervals: they guided the design, they do not prove it.

Counting stops being a guess

Counting questions answered correctly, all four models together (out of 24)

  • Turtle : 12/24
  • JSON-LD : 13/24
  • ISOLDA : 22/24
Exploratory pilot runs: 6 questions of the form “how many X are there?” × 4 models, same frozen questions and inference setup. In the pre-registered DCAT run the pattern held: ISOLDA got all 4 count questions right with every model; Turtle with labels got none.

04 · Where it fell short

The first real test: good on one graph, bad on the other

Saving tokens is not a success on its own, so ISOLDA was held to three criteria, fixed in advance: be lossless, be cheaper, and be understood at least as well as the best standard format. Version 0.1 was then tested on two graphs nobody had looked at during its design: the Catalan open data catalogue and the municipalities of Catalonia from Wikidata.

It was lossless and cheaper. But overall, it was significantly less accurate than Turtle annotated with labels, for all four models. The breakdown by graph shows where:

Overall accuracy in the pre-registered confirmatory run (higher is better)

Gemma 4 31B · DCAT

  • JSON-LD : 91 %
  • Turtle + labels : 91 %
  • ISOLDA 0.1 : 92 %

Gemma 4 31B · Municipalities

  • JSON-LD : 75 %
  • Turtle + labels : 79 %
  • ISOLDA 0.1 : 36 %

Qwen3.8 27B · DCAT

  • JSON-LD : 91 %
  • Turtle + labels : 90 %
  • ISOLDA 0.1 : 100 %

Qwen3.8 27B · Municipalities

  • JSON-LD : 72 %
  • Turtle + labels : 77 %
  • ISOLDA 0.1 : 52 %

Gemma 4 E4B · DCAT

  • JSON-LD : 72 %
  • Turtle + labels : 78 %
  • ISOLDA 0.1 : 94 %

Gemma 4 E4B · Municipalities

  • JSON-LD : 37 %
  • Turtle + labels : 65 %
  • ISOLDA 0.1 : 36 %

Qwen3 4B · DCAT

  • JSON-LD : 61 %
  • Turtle + labels : 64 %
  • ISOLDA 0.1 : 63 %

Qwen3 4B · Municipalities

  • JSON-LD : 26 %
  • Turtle + labels : 27 %
  • ISOLDA 0.1 : 14 %
87 questions on the DCAT catalogue and 170 on the municipalities graph, generated automatically from the structure of each graph and frozen before the run. “Turtle + labels” is Turtle with each linked entity's name added as a comment, the strongest baseline for every model.

On the data catalogue ISOLDA matched or beat the best baseline with every model — up to 94 % against 78 % for the small Gemma. On the municipalities it collapsed, most of all on multi-hop questions (2 % against 98 % for Gemma 4 31B).

Why. Every municipality in Wikidata has its name in Catalan, Spanish and English. Version 0.1 only used a name as the identifier when there was exactly one, so it fell back to slugs like manresa or bages — and the models answered with the slug, the very failure ISOLDA was built to avoid. The development graph, with one name per entity, never exposed it.

05 · What comes next

ISOLDA 0.2, and fresh data to test it

Version 0.2 changes four things, each aimed at a limit the tests exposed:

  1. A preferred language for names. With names in several languages, it picks one (English, then untagged, then the rest) as the identifier; the others become ordinary columns.
  2. Names row by row. One repeated name no longer turns a whole table into slugs; only that row does.
  3. Readable property names. When the graph knows what wdt:P1082 means, ISOLDA writes population.
  4. Shorter links. Entity IRIs are written with a prefix instead of in full on every row.

On the data used to design it, 0.2 already cuts the KG-LLM-Bench contexts from 8,076 to 5,925 tokens on average. That is not a validation, so two new pre-registered studies are running: one on a graph of writers in Catalan from Wikidata, and one on KG-LLM-Bench itself, with its own tasks, prompts and scoring. A third will ask whether the results hold when the model works entirely in Catalan, a non-hegemonic language.

This page will be updated as results come in, whichever way they go.