Draft Xavier Vinaixa Roselló
ISOLDA: Indexed Subject-Ordered Linked Data for AI
Draft. The Catalan version is kept in step with this one. First draft of the abstract and sections 1, 3 and 4, and of the final results in section 5; sections 6 and 7 are outlines. Everything will be revised when the validation of ISOLDA 0.2 and the KG-LLM-Bench study are analysed. All figures come from the project's lab notebook and the analyses of each run.
Abstract
Provisional. Written before the two studies still running have finished.
Knowledge graphs in RDF or JSON-LD can be given to a language model as text, usually in Turtle or JSON-LD. We present ISOLDA (Indexed Subject-Ordered Linked Data for AI), a text serialization of RDF graphs meant to be read by language models, including small ones. ISOLDA writes one table row per entity, grouped by type, uses each entity's own label as its identifier, writes the number of values of every list, and starts each document with a short glossary of the constructs it uses. Decoding an ISOLDA document gives back a graph isomorphic to the original. We evaluated ISOLDA 0.1 in a pre-registered study on two graphs not seen during its design (a DCAT data catalogue and a Wikidata extract of Catalan municipalities) with four open-weight models (Qwen3 4B, Gemma 4 E4B, Qwen3.8 27B and Gemma 4 31B). Every document decoded without loss, and on the DCAT graph ISOLDA used 20.3–28.6 % fewer tokens than Turtle annotated with entity labels. But its overall accuracy was significantly lower than that of Turtle with labels for all four models. The deficit came from the municipalities graph. ISOLDA 0.2 was designed after this result, and it is now being validated on fresh data, together with a study on the KG-LLM-Bench benchmark. Their results are not yet available.
1. Introduction
Knowledge graphs published as RDF or JSON-LD, such as the schema.org markup embedded in web pages, can be given to a language model as part of its prompt. The model then has to read the graph from a text serialization.
How a graph is written affects how well a model reasons about it. Fatemi, Halcrow and Perozzi (ICLR 2024, arXiv:2310.04560 (opens in a new window)) compared nine text encodings of graphs and found that the choice of encoding improved performance by between 4.8 % and 61.8 %, depending on the task. For finding the nodes connected to a given node (PaLM 62B, zero-shot), an encoding that puts all the neighbours of a node on one line reached 53.8 %, against 19.8 % for an adjacency list. KG-LLM-Bench (Markowitz, Galiya, Ver Steeg and Galstyan, 2025, arXiv:2504.07087 (opens in a new window)) evaluated five formats on subgraphs of a real knowledge graph derived from Wikidata. Average accuracy was 0.42 for structured JSON, 0.41 for a list of edges and for YAML, 0.35 for Turtle and 0.34 for JSON-LD. Turtle and JSON-LD were also the most expensive prompts (8,171 and 13,503 tokens on average, against 2,645 for the list of edges). The best format depended on the model. Models aggregated outgoing edges much better than incoming ones, and accuracy stayed above 50 % up to degree 4 and then fell towards 10 %. Loveland et al. (KDD 2026, arXiv:2606.15633 (opens in a new window)) propose a mechanism that explains part of this: with rotary positional embeddings, the attention between two nodes that are adjacent in the graph decays as they move apart in the sequence.
In KG-LLM-Bench, the Turtle and JSON-LD prompts name entities and relations with opaque identifiers (ex:R1, ex:101) and give their labels separately, while the list of edges and the JSON formats use the labels directly. The cheapest format, the list of edges, is not a lossless serialization of RDF. We wanted a format that keeps everything that makes RDF what it is (IRIs, datatypes, language tags, blank nodes) and still reads well.
ISOLDA has three success criteria, in this order: (1) RDF → ISOLDA → RDF must give an isomorphic graph; (2) ISOLDA must take fewer tokens than Turtle and JSON-LD; (3) small models must understand it at least as well as the best standard format. Saving tokens is not a success on its own.
We designed ISOLDA on a single development graph, with a series of small pilot runs, and then tested it in a pre-registered confirmatory run on two graphs that were not inspected during the design. ISOLDA 0.1 met criteria 1 and 2 there, but not criterion 3. We report this result in full. We then designed ISOLDA 0.2 to remove the limits that the confirmatory run exposed, and we are validating it on new data.
The contributions of this paper are:
- The ISOLDA notation and a reference Python library that encodes, decodes and checks it. The output is deterministic, byte for byte.
- A set of exploratory pilot runs that separate three effects: explicit counts help counting; opaque identifiers, including readable slugs, hurt multi-hop questions in small models; and using the entity's label as its identifier removes most of that penalty.
- A pre-registered confirmatory evaluation of ISOLDA 0.1 on two held-out graphs and four models, with a negative result on the comprehension criterion.
- ISOLDA 0.2 and two pre-registered studies that are still running: a validation on a new Wikidata graph of writers in Catalan, and a study with the tasks, instances, prompts and scoring of KG-LLM-Bench.
Section 2 discusses related work, Section 3 describes the notation, Section 4 the experimental setup, and Section 5 the results available so far.
3. The ISOLDA notation
The name spells out the design. Indexed: every entity has a readable id, usually its own name, and it is the key every link (>) points to. The document can also carry explicit indexes: @types gives how many entities of each type there are, [N] gives the number of values in every list, and ~in lists incoming links. The ids and [N] are always there; @types can be turned off, and ~in is off by default. Subject-Ordered: all the statements about an entity are together, one table row or block per subject, grouped by type. Linked Data: it is a serialization of RDF / JSON-LD, and decoding it gives back a graph isomorphic to the original. for AI: it is meant to be read by language models, including small ones, and each document starts with a short glossary of only the constructs it uses.
3.1 Document structure
An ISOLDA document has three parts: a glossary, a few directives, and one section per class. Listing 1 shows an excerpt of the development graph (Section 4.1) in ISOLDA 0.2.
Glossary. The document starts with lines beginning with # that explain the notation to the model in English. The glossary is generated, not fixed: it contains only the entries for the constructs that appear in that document. The wording of each entry is fixed and versioned, so the same set of constructs always gives the same glossary, byte for byte. Its tokens count in every measurement in this paper.
Directives. @vocab declares the vocabulary whose terms are written without a prefix, and @prefix declares a prefix for every namespace used three or more times. @types gives the number of entities of each type in the whole graph, in descending order.
Sections. Each entity goes in the section of its main class: of its types, the one with the most entities in the graph. Its other types go in an a: property, and untyped entities go in a section Thing_. A section is a table, Type[N]{columns}:, with one row per entity. N is the number of entities of the type. The columns are the properties present in at least half of the rows. The properties of an entity that are not columns go to a section ~extra: at the end of the document, written as blocks under the entity's id. When every row of a table has the same value for a property, that value moves to the header as a constant, p=v.
Tables are the default. A section can also be written as blocks: one entity per id: line, followed by one property per line. The choice depends on a regularity threshold θ; by default θ = 0, so every section is a table and every entity takes exactly one line in its section.
3.2 Identifiers and links
Every entity has a readable id. By default the id is the entity's label, taken from the first label predicate present (schema:name, schema:headline, rdfs:label, dct:title, foaf:name, skos:prefLabel), when no other entity has the same label. If every row of a table uses its label as id, the first column is named after that property ({name, …}): the cell is both the id and the value of the label. Otherwise the id is a slug made from the label, from the local name of the IRI, or from the type and a number (organization-role-1). The entity's IRI, if it has one, goes in an @iri column.
A link to another entity is written with that entity's id, and the > marker goes once in the column header or the property key (memberOf>, author> Xavier Vinaixa Roselló), not in every cell.
3.3 Values
Each column is a pair (predicate, value class): link, IRI, plain text, text with a language tag, or typed literal. The header carries the datatype or the language (copyrightYear:int, title@ca), so a cell never needs an exception. A value with several elements is written [N] a | b | c, always with its count. An empty cell means that the entity has no value for that column; "" is the empty string. Values are quoted, as JSON strings, only when they are empty, start or end with spaces, start with [, or contain a comma, a bar, a quote, a brace or a line break. IRIs are written with a declared prefix when possible.
3.4 Order and determinism
Entities are ordered by breadth-first search from the node with the most outgoing triples; sections follow the first appearance of their class, and rows follow the search order. Blank nodes are canonicalized before encoding (colour refinement with individualization, falling back to rdflib's canonicalization only if ties remain) and relabelled b1..bn. No order depends on the iteration order of rdflib or on Python's hash randomization: the output is identical byte for byte under different values of PYTHONHASHSEED.
3.5 Options
The defaults were chosen after the pilot runs (Section 5.1). Other options exist for ablation: slugs as ids for every entity, blank nodes written inline inside the entity that references them, blocks instead of tables, continuation lines under each row instead of ~extra, no constants, no @types, a one-line legend or no glossary, an inverse index (~in:, incoming links grouped by target, for targets with two or more), and an anchor mode that orders the document from the entity a question is about and can crop it to k hops (cropped documents are not lossless with respect to the full graph). The decoder ignores @types and ~in, which are redundant. Named graphs are written as @graph sections.
@prefix line, the end of the @types line, 16 sections and most of the ~extra section.# ISOLDA 0.2: an RDF graph as text. Entities are grouped by type; names without a prefix are terms of @vocab.
# `@types` gives how many entities of each type the graph has.
# `pfx:rest` is a full IRI written with a declared @prefix.
# `Type[N]{a, b, c}:` is a table of the N entities of Type: one row per entity, cells in column order, first cell = the entity id.
# `@iri` is the entity's full IRI.
# A property ending in `>` links to another entity: its value is that entity's id.
# `[N] a | b` means N values.
# An empty cell means the entity has no value for that column.
# `p:date`, `p:int`, `p:iri`, `p:<type>`: every value of p has that datatype.
# `p=v` after a table header: every entity in the table has value v for p.
# `~extra:` lists more properties of entities from the tables above, by id.
@vocab http://schema.org/
@prefix xaviviro: https://xaviviro.com/#
[…]
@types Article 26, Organization 11, Person 6, ScholarlyArticle 6, Language 3, Occupation 3, OrganizationRole 3, […]
[…]
Occupation[3]{name, skills}:
Science Communicator,
Chief Technology Officer, [5] Artificial Intelligence | Computer Vision | Machine Learning | Natural Language Processing | Software Engineering
AI Researcher, [5] Fine-tuning | Large Language Models | LoRA | Sign Language Technology | Speech Synthesis
[…]
OrganizationRole[3]{id, roleName, memberOf>}:
organization-role-1, Technical Lead,
organization-role-2, President, ABAT — Associació de Birra Artesana de Tiana
organization-role-3, President, Opendata.cat — Associació pel foment de les dades obertes
[…]
~extra:
[…]
Chief Technology Officer:
occupationLocation> "Barcelona, Spain"
occupationalCategory: 15-1252.00
[…] In Occupation, the first cell is the label and the id at the same time. The first occupation has no skills, so its second cell is empty. The two properties of "Chief Technology Officer" that are not columns are in ~extra. The roles have no label, so their ids are slugs and the first column is id. memberOf> links each role to an organization by that organization's name. A header with constants from the same document is Article[26]{headline, image, url} author>=Xavier Vinaixa Roselló, inLanguage=ca, …: the 26 articles share their author and language, which are written once.
3.6 Changes in ISOLDA 0.2
ISOLDA 0.1 is the version tested in the confirmatory run (Section 4.4). ISOLDA 0.2 changes four things, each designed after that run and the first KG-LLM-Bench study to remove a limit they exposed (Section 5.2):
- Preferred label language. 0.1 used a label as id only if the entity had exactly one label literal; with labels in Catalan, Spanish and English, it fell back to slugs. 0.2 picks a preferred label (English, then untagged, then the rest) and uses it as id if it is unique. Labels in other languages become ordinary columns.
- Label ids row by row. In 0.1, a single repeated label made the whole table use slugs. In 0.2, a table keeps its label column if at least half of its rows have their label as id; the other rows write
@followed by their slug in that cell, meaning "this is only the id". - Readable names for properties and classes. When the graph contains the label of a property or a class, 0.2 declares an alias with
@term population = wdt:P1082and writes the alias wherever 0.1 wrote the term. - Compact IRIs from the
@vocabnamespace. When entity IRIs are in the@vocabnamespace, 0.2 also declares it as a prefix and uses it for values, instead of writing the full IRI in every row.
The development graph does not use any of these constructs, so Listing 1 is the same in 0.1 except for the first line of the glossary (the whole document differs in only one other line, where 0.2 quotes a value that starts with @). Options(version="0.1") reproduces the output of 0.1 byte for byte; this was checked on the 528 frozen ISOLDA contexts of the earlier runs. On the development data (o200k tokens), the mean size of the pseudonymized KG-LLM-Bench contexts falls from 8,076 to 5,925 tokens (their rdf_turtle2: 7,263), the municipalities sample from 9,484 to 8,617, and the DCAT sample from 13,965 to 13,640. These figures come from data that was used to design 0.2; they are not a validation.
3.7 Implementation
The reference implementation is the Python package isolda, which depends only on rdflib (opens in a new window). It reads JSON-LD (including the <script type="application/ld+json"> blocks of an HTML page), Turtle, N-Triples, N-Quads, TriG and RDF/XML, and has a command-line tool: isolda encode, isolda decode and isolda check, which exits with 0 if the round trip is isomorphic. The tests check isomorphic round trips on the development graph and on synthetic edge cases (typed and untyped strings, language tags with regions, separators and quotes inside values, shared, cyclic and empty blank nodes, entities with several types or none, label and slug collisions, named graphs) under every combination of options, and identical output under three values of PYTHONHASHSEED.
4. Experiments
All the studies share a few rules. Every graph format is decoded back to triples and compared with the reference graph with rdflib.compare.isomorphic; losslessness is checked, never assumed. Blank nodes are canonicalized and questions and contexts are frozen, with their SHA-1 hashes, before a run. Answers are appended to a log that is never rewritten, and discarded runs are kept and documented in the lab notebook. Reasoning ("thinking") is disabled in every model, so we measure the direct reading of the format, and decoding uses temperature 0.
4.1 Graphs
| Graph | Source | Size | Role |
|---|---|---|---|
| Development | The 43 JSON-LD blocks (schema.org) of the author's website, xaviviro.com/ca, 2026-09-29 | 810 triples; 717 after merging blank nodes with identical content | Design of 0.1 and pilot runs |
| DCAT | Catalogue of the open data portal of the Generalitat de Catalunya, RDF/XML, 2026-09-29 | 39,842 triples, 919 datasets; sample of 1,292 triples | Confirmatory run (held out) |
| Municipalities | Wikidata, municipalities of Catalonia, CONSTRUCT query, 2026-09-29 | 21,031 triples, 947 municipalities; sample of 1,235 triples | Confirmatory run (held out) |
| KG-LLM-Bench Countries | Instances regenerated with the benchmark's own code (seeds 41–45) | 5 tasks × 100 instances, each with its own subgraph of at most 200 edges | KG-LLM-Bench study (pseudonymized); 0.2 validation (real names) |
| Writers | Wikidata, writers who write in Catalan, 2026-09-30 | 4,038 writers; 400 downloaded in an order shuffled with seed 1202 (7,456 triples); sample of 1,118 triples, 49 writers | 0.2 validation (held out) |
Merging blank nodes with identical content in the development graph is equivalent under RDF simple entailment in both directions. The held-out graphs were downloaded, frozen with their hashes, and cut down deterministically: the sample is the longest prefix of units (datasets, groups of municipalities, writers) whose largest context among all the formats fits in 20,000 o200k tokens. They were not inspected: the scripts printed only sizes and hashes.
4.2 Models and inference
The first pilot run used LM Studio on an Apple M3 Max (36 GB) with MLX models: Qwen3 4B at 4 bits and Bonsai 27B at 1 bit. From the second pilot run on, all runs use vLLM (opens in a new window) on a rented GPU (RTX PRO 6000, 96 GB) with Qwen3 4B, Gemma 4 E4B and Gemma 4 31B in bf16, and Qwen3.8 27B in FP8 (the bf16 version entered a restart loop on the pod). The KG-LLM-Bench studies add Llama 3.2 1B (bf16) and Llama 3.3 70B (FP8). Tokens are counted with each model's own tokenizer, except where o200k is stated. Figures from the two inference environments are not comparable.
4.3 Pilot runs (exploratory)
The pilot runs use the development graph and 35 questions generated automatically from it, each with a known answer: 15 lookups (a literal property of an entity named in the question), 14 multi-hop questions (the name of the entity reached by a relation) and 6 counts (the number of entities of a type). The same frozen questions are used in all five pilot runs. The data goes in the system message and the question in the user message. Answers are scored automatically after NFKC normalization and case folding: for counts, the first integer in the answer; otherwise, equality or the expected answer as a whole phrase inside a short answer.
- run1: JSON-LD (the original blocks), Turtle, and two early prototypes with one table per type. The lab notebook calls them TOON-LD v1 and v2 because they were built with a TOON encoder (Section 2). v1 refers to entities by opaque ids (
@p3); v2 uses entity names as keys. The run was stopped early because the local server stopped reusing its prompt cache (about 150 s per question instead of 5 s). Its figures combine the frozen run with answers recovered from the logs of a preliminary run on the same questions, whose Turtle and prototype contexts had the same content but a different blank node numbering and column order. v2 was not evaluated. - run2: the same inputs, with the four models on vLLM.
- run3 and run4: the first ISOLDA prototype, with slugs as ids (run3) or labels as ids (run4), each with and without inline blank nodes.
- run5: four variants of the run4 prototype with labels as ids and no inline nodes: the reference, with
@types, with one line per entity, and without constants.
The prototype v2 and the ISOLDA prototypes were designed after seeing the results of run1 on the same questions. The pilots are therefore exploratory: they guided the design and fixed the token threshold of the confirmatory run, and they can overfit the development graph.
4.4 Confirmatory run (pre-registered)
The confirmatory run tests ISOLDA 0.1, with the defaults chosen after run5, on the DCAT and municipalities samples. Everything below was fixed and recorded in the lab notebook before the run, and the questions and contexts were not read.
Questions. A generic question generator, driven only by the structure of the graph, with seeds 1101 (DCAT) and 1102 (municipalities), and at most 40 questions per family and graph:
| Family | DCAT | Municipalities |
|---|---|---|
| lookup | 40 | 40 |
| multi-hop (1 hop) | 0 | 40 |
| multi-hop (2 hops) | 1 | 0 |
| count by type | 4 | 1 |
| outgoing aggregation (degree 1, 2, 4, 8, 16) | 40 | 40 |
| incoming aggregation | 2 | 40 |
| filter on typed literals | 0 | 9 |
| language choice | 0 | — |
| Total | 87 | 170 |
The DCAT graph gave few structural questions. Having seen only these counts, we decided to keep the generator unchanged.
Formats (arms). Three baselines: B-json, compact JSON-LD with the nodes referenced only once nested inside their referrer; B-ttl, Turtle; and B-ttl-ids, Turtle with the label of each linked entity in a comment at the end of the line. B-ttl-ids is the lossless form of "Turtle with readable ids", since IRIs cannot be renamed. All three use the same @vocab and prefixes as ISOLDA. The main arm, I, is ISOLDA 0.1 with its defaults. Eleven ablations change one thing each: slugs as ids, inline blank nodes, the hybrid table/block layout of run4, no @types, continuation lines instead of ~extra, blocks only, no glossary, a one-line legend, the inverse index, and the anchor mode. Qwen3 4B and Gemma 4 E4B read all 14 arms; Qwen3.8 27B and Gemma 4 31B read B-json, B-ttl-ids and I. max_tokens was 200.
Success criteria. They come from the experimental plan written before the pilots, with the token threshold fixed after the pilots:
- 100 % of the round trips are isomorphic.
- I takes at least 15 % fewer tokens than B-ttl-ids on DCAT, with each model's tokenizer.
- For no model is the overall accuracy of I significantly lower than that of the best baseline.
- I is significantly better than the best baseline in at least one family: aggregation with degree above 4, incoming aggregation, or multi-hop with the 4B models.
If only criteria 1 and 2 were met, the result would be reported as "cheaper and as good", not as an improvement in reasoning.
Statistics. Exact two-sided McNemar tests, α = 0.05. For criteria 3 and 4, one test per model with the questions of both graphs pooled (models are never pooled); the best baseline of each model is the one with the highest overall accuracy on both graphs. For criterion 4, one test per family and model. There is no correction for multiple comparisons; all p-values are reported. 95 % confidence intervals come from a bootstrap with 10,000 resamples.
4.5 KG-LLM-Bench study (pre-registered, in progress)
This study runs ISOLDA 0.1, unchanged, with the tasks, instances, prompt and scoring of KG-LLM-Bench. It has two aims: to reproduce the published figures of the benchmark's two open models, which validates our setup, and to compare ISOLDA with the benchmark's formats under the same conditions.
Instances. The original instances are in a remote store that is not public, so we regenerated them with the benchmark's own script and configuration (seeds 41–45). Generation is deterministic: the 500 instances (five tasks × 100) came out identical under different hash seeds and library versions. We use the pseudonymized variant. The correct answer scores 1 in 494 of the 500 instances; the other six are HighestDegreeNode instances whose answer contains a hyphen or an apostrophe, which the benchmark's answer parser does not accept, so no format can get them right.
Formats. The five formats of the paper (list_of_edges, structured_yaml, structured_json, rdf_turtle3, json_ld3; the last two use opaque ids), two variants that the benchmark's code can also produce (rdf_turtle2, json_ld2: entity ids with rdfs:label and readable relation IRIs), and isolda, encoded from the same RDF as rdf_turtle2 (isomorphic in all 500 instances). The prompt is the benchmark's, as a single user message; temperature 0 and max_tokens 512, as in the benchmark; the scoring is the benchmark's own function, unchanged.
Models. Llama 3.2 1B, Qwen3 4B and Gemma 4 E4B read all eight formats; Llama 3.3 70B reads the paper's five, rdf_turtle2 and isolda; Qwen3.8 27B and Gemma 4 31B read list_of_edges, structured_json, rdf_turtle2 and isolda. In total, 19,500 calls.
Criteria (α = 0.05, exact McNemar paired by instance):
- R. For Llama 3.2 1B and Llama 3.3 70B, in each of the paper's five formats, our overall accuracy is within ±0.10 of the published one.
- K1. For each model, the mean token count of ISOLDA is lower than that of
rdf_turtle3,json_ld3,rdf_turtle2andjson_ld2. - K2. For no model is the overall accuracy of ISOLDA significantly lower than that of the best of the paper's formats that the model ran.
- K3. For each model that ran
rdf_turtle3, ISOLDA is significantly more accurate thanrdf_turtle3.
Before any model was run, we measured the contexts of one task with o200k: ISOLDA 7,968 tokens, rdf_turtle3 6,926 and rdf_turtle2 7,103. K1 was therefore expected to fail. We traced the cost to two limits of 0.1 (full IRIs in every row, and slugs for a whole table when one label is repeated), but left the library unchanged for this study, since changing it after seeing these data would be overfitting. Both limits are addressed in 0.2. The study was running at the time of writing.
4.6 Validation of ISOLDA 0.2 (pre-registered, in progress)
ISOLDA 0.2 was designed after the confirmatory run and the first KG-LLM-Bench study, so the DCAT, municipalities and pseudonymized KG-LLM-Bench data are now development data. The validation uses two other data sets.
Writers. People (P31 = Q5) who are writers (P106 = Q36180) and write in Catalan (P1412 = Q7026) in Wikidata (opens in a new window): dates and places of birth and death, occupations, works with their type and date, awards and languages, with Catalan, Spanish and English labels for everything linked and English labels for the properties. The graph is heterogeneous, multilingual and uses Wikidata property codes, so it tests changes 1–3. The question generator of the confirmatory run, with seed 1201 and all eight families, gave 201 questions: 40 each of lookup, multi-hop, outgoing aggregation, incoming aggregation and language choice, and one count; it found no candidates for two-hop or filter questions in this graph. Arms: B-json, B-ttl, B-ttl-ids, ISOLDA 0.2, ISOLDA 0.1, and 0.2 without each of its four changes. All ISOLDA arms are presented to the model with the same name. Qwen3 4B and Gemma 4 E4B read the nine arms; Qwen3.8 27B and Gemma 4 31B read B-json, B-ttl-ids, 0.1 and 0.2.
KG-LLM-Bench with real names. The same 500 instances without pseudonyms, in the benchmark's five formats, rdf_turtle2, ISOLDA 0.1 and ISOLDA 0.2. Llama 3.2 1B, Qwen3 4B and Gemma 4 E4B read the eight formats; Qwen3.8 27B and Gemma 4 31B read list_of_edges, structured_json, rdf_turtle2 and both ISOLDA versions. This is a known weakness: it is the same Countries graph that was used, pseudonymized, during development.
Criteria (α = 0.05, exact McNemar per model, no correction for multiple comparisons):
- W1–W4 on the writers graph: the four criteria of the confirmatory run, applied to ISOLDA 0.2 (lossless; at least 15 % fewer tokens than B-ttl-ids for every model; never significantly worse than the best baseline; significantly better in at least one family).
- K1–K3 on KG-LLM-Bench, applied to ISOLDA 0.2 (for K1, against
rdf_turtle3,json_ld3andrdf_turtle2). - V on each data set: ISOLDA 0.2 is significantly more accurate than ISOLDA 0.1 for at least two of the four main models, and significantly less accurate for none of the models that ran both.
- Exploratory: each ablation against 0.2, by family and model.
All nine writers contexts and both ISOLDA versions of the 500 KG-LLM-Bench instances decode without loss. After the runs had been frozen and started, and before any result was seen, the author looked at the beginning of the ISOLDA context of the writers graph. Nothing in these runs was changed; anything learned from it can only inform a later version, which would need another graph.
4.7 KG-LLM-Bench in Catalan (planned)
The third study asks whether the results hold when the model works in Catalan, a non-hegemonic language. It uses the same 500 KG-LLM-Bench instances in two variants: real names, with the Catalan Wikidata label of each entity and relation, and the paper's pseudonyms, with relations, instructions and questions in Catalan. Everything the model reads is in Catalan except the ISOLDA glossary, which stays in English. Entities without a Catalan label (about 30 % of the 7,079 entities in the instances, 10 % per instance at the median) keep their original name. The answer is asked as «Resposta: …», and the paper's scorer is adapted in two ways only: it accepts the Catalan keyword, and it accepts apostrophes, hyphens and the middle dot in entity labels (589 Catalan labels have one of them). The English runs are rescored with the same scorer. The models are the four modern ones; Llama 3.2 is left out because it does not officially cover Catalan. The criteria are pre-registered: C1 lossless, C2 fewer tokens than every RDF format, C3 never significantly worse than the best paper format, C4 better than rdf_turtle3, and C5, specific to this study: ISOLDA's drop from English to Catalan is never significantly larger than that of the best paper format. A criterion passes only if it holds in both variants. The ISOLDA version will be fixed when the validation of ISOLDA 0.2 ends.
5. Results
5.1 Pilot runs (exploratory)
All the results in this subsection come from the development graph and the same 35 questions, with small n and no confidence intervals. They are exploratory. Lookups were answered correctly by every model in every format and run (15/15), so the tables show the other families.
run1 (LM Studio, MLX; correct answers out of 35, then multi-hop /14 and count /6):
| Model | JSON-LD | Turtle | Prototype v1 (opaque ids) |
|---|---|---|---|
| Bonsai 27B (1 bit) | 30 (13 · 2) | 27 (10 · 2) | 27 (6 · 6) |
| Qwen3 4B (4 bits) | 30 (13 · 2) | 23 (7 · 1) | 23 (2 · 6) |
With opaque ids the models answered the id (@p3, @o15) instead of the name it points to; with Turtle they answered blank node labels or IRIs. JSON-LD, which nests the referenced object with its name, did best on multi-hop. The prototype, which writes [N] in the header, answered every count correctly.
run2 (vLLM; correct /35, multi-hop /14, count /6, context tokens with each model's tokenizer):
| Model | JSON-LD | Turtle | Prototype v1 | Prototype v2 (names as keys) |
|---|---|---|---|---|
| Qwen3 4B | 29 · 11 · 3 · 12,028 | 25 · 8 · 2 · 12,118 | 24 · 3 · 6 · 10,754 | 33 · 12 · 6 · 13,875 |
| Gemma 4 E4B | 31 · 12 · 4 · 12,791 | 29 · 12 · 2 · 13,354 | 29 · 8 · 6 · 11,579 | 33 · 13 · 5 · 14,585 |
| Qwen3.8 27B (FP8) | 32 · 14 · 3 · 11,893 | 33 · 14 · 4 · 12,664 | 35 · 14 · 6 · 10,725 | 35 · 14 · 6 · 13,690 |
| Gemma 4 31B | 32 · 14 · 3 · 12,795 | 33 · 14 · 4 · 13,358 | 35 · 14 · 6 · 11,583 | 35 · 14 · 6 · 14,589 |
Explicit counts won on all four models: the prototypes got 5–6 of 6 counts, JSON-LD and Turtle 2–4, including the 27–31B models. Opaque ids hurt only the small models (multi-hop 3/14 and 8/14, against 14/14 for the large ones). Prototype v2 was the best format for every model, at the cost of about 15 % more tokens than JSON-LD; it was designed after run1, so this comparison is not clean.
run3 and run4 (first ISOLDA prototype; multi-hop /14 · count /6 · total /35):
| Model | JSON-LD | Prototype v2 | Slugs | Slugs, no inline | Labels | Labels, no inline |
|---|---|---|---|---|---|---|
| Qwen3 4B | 11 · 3 · 29 | 12 · 6 · 33 | 8 · 4 · 27 | 5 · 6 · 26 | 11 · 4 · 30 | 12 · 6 · 33 |
| Gemma 4 E4B | 12 · 4 · 31 | 13 · 5 · 33 | 10 · 2 · 27 | 3 · 3 · 21 | 13 · 3 · 31 | 13 · 4 · 32 |
| Qwen3.8 27B (FP8) | 14 · 3 · 32 | 14 · 6 · 35 | 13 · 5 · 33 | 13 · 6 · 34 | 14 · 5 · 34 | 14 · 6 · 35 |
| Gemma 4 31B | 14 · 3 · 32 | 14 · 6 · 35 | 13 · 4 · 32 | 13 · 6 · 34 | 14 · 4 · 33 | 14 · 6 · 35 |
Readable slugs did not solve the indirection: the small models answered with the slug (catalunya-radio, sorensen-ai). With the label as id, they reached 12/14 and 13/14 on multi-hop. ISOLDA with labels and without inline nodes matched prototype v2 (135/140 against 136/140) with about 28 % fewer tokens (10,042 against 13,875 with the Qwen3 4B tokenizer). Inline blank nodes hurt counting through a design error: Person[4] in a header when the type had six entities, two of them nested.
run5 (variants of ISOLDA with labels as ids; correct /35):
| Model | Reference | + @types | One line per entity | No constants |
|---|---|---|---|---|
| Qwen3 4B | 33 | 34 | 35 | 34 |
| Gemma 4 E4B | 32 | 33 | 32 | 29 |
| Qwen3.8 27B (FP8) | 35 | 35 | 35 | 35 |
| Gemma 4 31B | 35 | 35 | 35 | 35 |
| Total /140 | 135 | 137 | 137 | 133 |
| Tokens (Qwen3 4B) | 10,042 | 10,197 | 9,613 | 10,964 |
The reference reproduced run4 exactly. @types fixed the counts of Gemma 4 E4B (4/6 → 6/6) for 1.5 % more tokens; one line per entity gave 35/35 with Qwen3 4B and 4 % fewer tokens; removing constants made Gemma 4 E4B worse and cost 9 % more tokens. These differences are of one or two questions and within noise. The defaults of ISOLDA 0.1 were fixed from them: labels as ids, no inline nodes, constants, @types, and one line per entity. The combination of @types and one line per entity was not tested in the pilots. With these defaults, the development graph takes 8,702 o200k tokens (Turtle: 11,058).
5.2 Confirmatory run (pre-registered)
The run produced all the expected answers (2,958 for DCAT and 5,780 for municipalities), with no errors and no reasoning tokens, and the frozen inputs matched their hashes.
| Criterion | Result |
|---|---|
| 1. Lossless | Pass. The 14 contexts of each graph, including every anchor-mode context, decode to an isomorphic graph. |
| 2. ≥ 15 % fewer tokens than B-ttl-ids on DCAT | Pass. 20.3 % (Qwen3 4B), 25.9 % (Qwen3.8 27B), 28.6 % (Gemma 4 E4B and Gemma 4 31B). |
| 3. Never significantly worse than the best baseline | Fail. For all four models the best baseline was B-ttl-ids, and ISOLDA was significantly worse (below). |
| 4. Significantly better in at least one family | Pass, through a single test (below). |
For criterion 3 (257 questions per model), the discordant pairs were, ISOLDA right only against B-ttl-ids right only: Gemma 4 31B 10 against 82 (95 % CI of the difference in accuracy −0.35 to −0.21); Gemma 4 E4B 31 against 67 (−0.21 to −0.06); Qwen3 4B 22 against 45 (−0.15 to −0.03); Qwen3.8 27B 11 against 45 (−0.19 to −0.08). All p < 0.01.
Criterion 4 passed only because of outgoing aggregation with degree above 4 on Gemma 4 E4B (15 against 1, p = 0.0005). In the same family, Gemma 4 31B went the other way (3 against 13, p = 0.02), and on multi-hop Gemma 4 E4B was right with B-ttl-ids and wrong with ISOLDA in 34 questions, and never the reverse. No test on incoming aggregation was significant.
The pre-registered conclusion is that ISOLDA 0.1 does not meet criterion 3. It cannot be presented as "as good as" the standard formats, only as lossless and cheaper.
Overall accuracy by graph shows where the deficit comes from:
| Model | DCAT | Municipalities | ||||
|---|---|---|---|---|---|---|
| JSON-LD | Turtle + labels | ISOLDA | JSON-LD | Turtle + labels | ISOLDA | |
| Gemma 4 31B | 91 % | 91 % | 92 % | 75 % | 79 % | 36 % |
| Qwen3.8 27B (FP8) | 91 % | 90 % | 100 % | 72 % | 77 % | 52 % |
| Gemma 4 E4B | 72 % | 78 % | 94 % | 37 % | 65 % | 36 % |
| Qwen3 4B | 61 % | 64 % | 63 % | 26 % | 27 % | 14 % |
On DCAT (87 questions), ISOLDA was within one point of Turtle with labels or above it for every model. On the municipalities (170 questions), it was far behind for every model. Multi-hop questions on the municipalities show the gap most clearly: ISOLDA 2 % against 98 % with Turtle and labels on Gemma 4 31B, 2 % against 100 % on Qwen3.8 27B, and 5 % against 90 % on Gemma 4 E4B.
Exploratory observations (not pre-registered; the ablations had no criterion):
- On DCAT, ISOLDA answered all four count questions correctly with every model; Turtle with labels answered none.
- The blocks-only ablation was the best arm on DCAT for the two small models (Qwen3 4B 85 %, Gemma 4 E4B 98 %). On the municipalities it was better than default ISOLDA (24 % and 51 %) but still below Turtle with labels.
- On the municipalities, the slug ablation gave exactly the same results as default ISOLDA for both small models.
Diagnosis (made after seeing the results). We found two limits of ISOLDA 0.1 that the development graph did not have. First, each municipality has labels in Catalan, Spanish and English; 0.1 used a label as id only when there was exactly one, so it fell back to slugs (manresa, bages), and the models answered with the slug, the same failure as in run3. The identical results of the slug ablation support this. Second, ISOLDA wrote Wikidata properties by their codes (P1082, P131). Our first reading was that Turtle with labels showed the property labels and ISOLDA did not. A later check found that the sampling code kept the labels of IRIs in object position only, so the property labels were missing from the municipalities sample for every format, and the questions named the properties by their codes. The second explanation therefore needs to be re-examined. ISOLDA 0.2 (Section 3.6) addresses both limits, and the two found in the KG-LLM-Bench setup, and it is being validated on new data.
5.3 Studies in progress
Outline.
- KG-LLM-Bench with ISOLDA 0.1 (Section 4.5): criteria R and K1–K3; results by task;
rdf_turtle2againstrdf_turtle3as a measure of the effect of opaque ids in the benchmark's own format. - Validation of ISOLDA 0.2 (Section 4.6): criteria W1–W4, K1–K3 and V; the four ablations by family and model.
- Token counts for every format, graph and tokenizer.
6. Discussion and limitations
Outline.
- Token savings alone are not enough: the pilots and the confirmatory run show a cheaper format that is worse for small models.
- Identifiers: the most consistent finding across the pilots and the confirmatory run is that small models answer with whatever id they see (opaque ids, blank node labels, slugs). Labels as ids address this (principle P1, hypothesis H1 of the plan); 0.2 extends them to multilingual graphs.
- Explicit counts (
[N],@types) and counting questions (hypothesis H3), with the caveat of very few count questions in the held-out graphs. - Tables against blocks: blocks did better on DCAT for small models, while the pilots favoured one line per entity. The evidence is contradictory; 0.2 keeps tables.
- Limits of the evidence: one development graph and one fixed question set for all the pilots; held-out graphs that are fairly regular; few structural questions on DCAT; the missing property labels in the municipalities sample; automatically generated questions; one run at temperature 0; reasoning disabled; no correction for multiple comparisons; changes of inference environment (MLX, vLLM, FP8).
- KG-LLM-Bench: regenerated instances, differences in inference with the published study (criterion R), and the reuse of the Countries graph in the 0.2 validation.
7. Conclusion
Outline.
- Summary of what the pre-registered studies support, and what they do not.
- The status of ISOLDA 0.2, depending on the validation.
- Future work: a study of KG-LLM-Bench in Catalan, with modern models, planned but not yet designed; runs with reasoning enabled; more irregular graphs; language models writing ISOLDA.