EFFKG Review, Part 3: How to Score Knowledge Graph Referencing Quality

This is Part 3 of the EFFKG Review. Part 1 imported the European Fishing Fleet Knowledge Graph (EFFKG) and identified 3,708,013 statement nodes associated with 13 distinct reference nodes. Part 2 examined the structure of these relationships: a vessel is associated with multiple statements, each statement is associated with a reference, and each reference identifies a registry. The resulting coverage is 99.9886%, which indicates that provenance information is available for almost all statements in the knowledge graph.
This part examines what a coverage value of this magnitude actually guarantees. The analysis applies a published framework rather than relying on an informal assessment. Four RQSS metrics are evaluated end to end, and the scores are reported according to the definitions provided by the framework, including cases where the resulting values are counter-intuitive. The analysis also reveals limitations in the framework itself that are not apparent from the quality of the dataset alone.
Four million provenance references in the European Fishing Fleet Knowledge Graph resolve to thirteen reference nodes. An examination of six vessels is sufficient to demonstrate the extent of this concentration.

What RQSS Measures
Referencing quality is a distinct and formalizable property of a knowledge graph and should be considered separately from data quality in general. A knowledge graph may contain accurate values while providing insufficient information to verify them. Conversely, it may provide references for all statements without providing enough information for users to verify those references.
RQSS formalizes this distinction through 40 metrics organized into 22 Linked Data quality dimensions and six categories. The framework specializes the quality taxonomy of Zaveri et al. to references rather than to facts. Of the 40 metrics, 34 are objective, meaning that they can be computed without human expert judgment. These 34 metrics are therefore implemented by the accompanying tool. When applied to subsets of Wikidata, the original study reports an average referencing quality score of 0.58 out of 1.
Hosseini Beghaeiraveri SA, Gray A, McNeill F. RQSS: Referencing quality scoring system for Wikidata. Semantic Web: Interoperability, Usability, Applicability. 2024;15(6):2419 to 2475. doi:10.3233/SW-243695. Open-access copy: swj3593.pdf. Implementation: RQSSFramework, CC0-1.0.
I was the lead author of the paper in which the RQSS framework was introduced, co-authored with my PhD advisors. The paper constitutes the main contribution of my PhD thesis, which represents a potential conflict of interest that should be stated explicitly. To address this, every result reported below is accompanied by the query or command used to obtain it, and the framework is subjected to at least as critical an evaluation as the dataset itself.
RQSS was developed for Wikidata, but its implementation can also be used for other datasets based on the Wikibase platform. The EFFKG is a separate Wikibase instance rather than a subset of Wikidata. The metrics can be transferred to any Wikibase deployment that exposes the relevant reification types because the queries use wikibase:Statement and wikibase:Reference in the http://wikiba.se/ontology# namespace. These terms are part of the Wikibase platform ontology and are not specific to the Wikidata instance.
The EFFKG query service does not expose these types. As reported in Part 1, querying ?s a wikibase:Statement and ?r a wikibase:Reference on the EFFKG endpoint returned zero results, whereas the corresponding queries on the imported copy returned 3,708,013 statement nodes and 13 reference nodes. Consequently, a dataset published for citation cannot be assessed for these metrics through the endpoint on which it is published. This is the first finding of this analysis and provides the motivation for the import described in Part 1. All subsequent analyses are performed on AWS Neptune. To the best of our knowledge, this is the first application of RQSS to a Wikibase-derived dataset other than Wikidata.
One limitation of the experimental setup should also be stated. The Neptune cluster used for the analysis is a db.t3.medium instance (two vCPUs and 4 GiB of memory) and contains 31.7 million triples. Two of the 25 reporting queries failed after three attempts each, and one aggregate query was not executed because of the observed resource limitations. These failures reflect the instance configuration used for this analysis and should not be attributed to Neptune, SPARQL, or the EFFKG itself. The instance was configured as described in the setup procedure. The four metrics below, and every reporting query behind them, run from run-all.sh in the EFFKG Review repository, which also holds the CSVs the figures are drawn from.
Metric 1: Availability of External URIs
This metric belongs to the Accessibility category, availability dimension. The score is the proportion of distinct external URIs that return HTTP 200. They are extracted from the reference nodes using a single query, excluding Wikibase and Wikidata namespaces. The framework’s -eu extractor exceeded the instance’s resource limits, so the equivalent scoped query was used instead, with its output passed unchanged to the runner:
PREFIX wikibase: <http://wikiba.se/ontology#>
SELECT DISTINCT ?to_ret WHERE {
?ref a wikibase:Reference ; ?refProperty ?to_ret .
FILTER (isIRI(?to_ret) &&
!CONTAINS(lcase(str(?to_ret)), "wikidata.org") &&
!CONTAINS(lcase(str(?to_ret)), "wikiba.se"))
}
| External URI | HTTP 200 |
|---|---|
webgate.ec.europa.eu/fleet-europa |
yes |
servicio.pesca.mapama.es/CENSO/ConsultaBuqueRegistro/Buques/Search |
yes |
www.mapa.gob.es/es/pesca/temas/registro-flota/informacion-sobre-flota-pesquera |
yes |
www.fao.org/fishery/en/area |
yes |
www.boe.es/eli/es/rd/2007/05/18/638/con |
yes |
boe.es/buscar/pdf/2007/BOE-A-2007-10951-consolidado.pdf |
yes |
signa.ign.es/signa/ |
yes |
msi.nga.mil |
yes |
app.hubocean.earth/catalog |
yes |
www.data.gouv.fr/api/1/datasets/r/60fe965d-5888-493b-9321-24bc3b1f84db |
yes |
openknowledge.fao.org/items/4de1cb7a-34e3-4645-ae94-0cfc3ef4ad70 |
no |
openknowledge.fao.org/handle/20.500.14283/cb5201en |
no |
openknowledge.fao.org/server/api/core/bitstreams/acee638a-d119-4fc6-8512-d7d953b14cd8/content |
no |
All three failures occurred on the FAO’s openknowledge repository. Because the sample contains only 13 URIs, one host changing its response can alter the score substantially. The score is therefore a point-in-time measurement, although this temporal limitation is not reflected in the output. An identical run the previous day returned 9 of 13 URIs and a score of 0.692. The URI that changed was the data.gouv.fr API link, which timed out while reaching static.data.gouv.fr in the first run but responded successfully in the second. The dataset itself did not change between the runs.
The metric measures whether a URI returns HTTP 200; it does not establish that the cited content remains available or unchanged. Several of the 13 URIs are search interfaces or API endpoints whose content can change over time. For example, servicio.pesca.mapama.es/CENSO/…/Search is a query form, while data.gouv.fr/api/1/datasets/r/… identifies a dataset resource that can be republished at the same URI. Neither therefore guarantees retrieval of the content as it existed when the graph was created. Assessing the content of references is the subject of other metrics. Nevertheless, the authors’ use of resolvable registry URLs rather than opaque internal identifiers is a deliberate and useful design choice.
Metric 12: Ratio of Reference Sharing
This metric belongs to the Intrinsic category, the conciseness dimension. It measures the proportion of reference nodes with more than one incoming prov:wasDerivedFrom edge. All 13 reference nodes meet this condition, giving 1.0, the maximum possible score.

13,13,1.0 for reference sharing and 1,13,0.923 for reference property diversity, followed by the per-node incoming counts.The per-node counts provide a more informative view:

prov:wasDerivedFrom edges per reference node, log scale. Two nodes carry 99.6% of the 4,034,214 references and the least used carries four. Source: RQSS Metric 12 (ref_sharing.csv) against the Neptune import, 9 August 2026.The 13 counts sum to 4,034,214. This is the only completed derivation of the total on this instance. A COUNT over more than four million wasDerivedFrom edges exhausts the cluster, while the per-source breakdown query failed three times. There is therefore no independent second derivation of this figure. The result is close to the 4,033,845 references reported by the authors for version 1.0.0, a difference of 369, comparable to the copy divergence measured in Part 1 between the endpoint and the distribution.
Two reference nodes account for 99.6% of all references. The largest accounts for 85.5% and the second for 14.1%, leaving 14,948 references across the remaining 11 nodes. The mapping between these nodes and their source URLs is not established because the required join query failed.
Wikibase content-hashes reference nodes, so identical reference content is represented by a single node. The 13 nodes therefore represent the distinct reference content used across the graph, rather than a sample or summary of it.
A perfect score on a conciseness metric should therefore be interpreted in the context of the metric’s assumptions. RQSS was designed for Wikidata, where extensive reference sharing can indicate contributors using the same convenient citation. In a bulk-integrated graph, however, extensive sharing can have a different interpretation: using one reference for each source registry can be an appropriate representation when the data are ingested from a small number of authoritative registries. The score of 1.0 therefore does not, by itself, indicate a problem with the dataset.
A narrower limitation is evident independently of the metric. A reference in this graph identifies the registry but not the specific record, file version, row, or retrieval. The provenance layer can therefore identify the registry from which information was obtained, but cannot by itself establish which version of that information is current when sources disagree.
Metric 32: Diversity of Reference Properties
This metric belongs to the Representational category and the representational-consistency dimension. It is defined as 1 - (reference properties / reference triples). The 13 reference nodes contain 13 reference triples using one distinct reference property, so 1 - 1/13 gives 0.923. Each reference node therefore contains exactly one triple: pr:P46 source. This is the complete provenance vocabulary used by the graph.
The fact that no reference node contains a date, affects Metric 17, freshness of reference triples, in the Dynamicity category. This metric requires the time since the last update relative to the total existence duration of the reference triples. Without a date on the reference node, the metric is uncomputable by definition. The graph already contains suitable date properties. Part 2 identified P44 date, P47 start date, and P48 end date, with 5,180, 130, and 57 uses respectively, all as qualifiers. Wikibase also supports these properties on references, as demonstrated by Wikidata’s use of retrieved (P813) with reference URL (P854). Adding a retrieval date to each reference would require one additional triple per reference and would make Metric 17 computable. This is therefore a high-value change for a future release.

P44 date, P47 start date and P48 end date are already defined in the graph but appear only in the qualifier position, so Metric 17 is the one metric a date on the reference would unblock. Qualifier counts from Part 2.Metric 14: Human-Added References
This metric belongs to the Trust category and believability dimension. The paper defines it as the ratio of human-added reference triples to all reference triples. In the implementation, RQSS evaluates referenced facts using item-page history, classifying an edit as bot-made when bot token appears in the username. However, not all datasets, including EFFKG, use bot token for declaring automated additions. Before this work, RQSS had only run on Wikidata but for this study, we fixed this limitation in commit f8f6c70.
Because the instance could not process all 190,649 EFFKG items, every 600th item was sampled, yielding 317 items and 5,718 referenced facts. The history of each sampled item was then retrieved and the first result was as follow:
num of items,num of referenced facts,num of human-added refed facts,num of not found facts,score
317,5718,5584,134,1.0
The score is 1.0, but it requires qualification. First, 134 of the 5,718 referenced facts had no retrievable history and were therefore excluded from the denominator, leaving 5,584 facts in the calculation. Second, the result depends on how the importing account is classified. The EFFKG import was performed by Admin, whose username does not contain bot and declaring this account as a bot changes the score to 0.0:
num of items,num of referenced facts,num of human-added refed facts,num of not found facts,score
317,5718,0,134,0.0
As shown above, depending on how Admin is classified, the same graph produce a score of 1.0 or 0.0. The available history does not establish whether Admin represents a bot, a human-operated bulk import, or a manual import. Hence, the metric cannot distinguish whether EFFKG follows Wikidata’s practice of explicitly identifying bot accounts, and the resulting score should be interpreted with caution.
Anyway, automated integration itself cannot be interpreted negative or positive quality indicator. This result should not be interpreted as a criticism of the dataset. Automated construction is inherent to large-scale data integration and provides a level of consistency that manual curation cannot readily achieve and WESO team do not claim otherwise. The relevant unresolved issue is whether the references actually support the statements to which they are attached. This is addressed by RQSS Metric 29 (relevance of reference triples), which requires subjective assessment and was not evaluated here.
When One Authority Contradicts Itself
As mentioned in Part 2, multiple references for the same property provide corroboration. That post found 326,623 of 3,707,591 referenced statements with more than one reference node. Multiple references may support the same value or different values; in the latter case, the graph preserves the disagreement between sources rather than resolving it. Both are expected in provenance-aware knowledge bases and are consistent with the provenance model used by standards such as PROV-O.
The case of interest however, is when a single reference authority supports two different values for the same property of the same entity. For example, on Q1000 (NUEVO SANTA MARIA), main engine power is recorded as 176.52 with the EU Fleet Register and the Spanish national registry as references, and separately as 240 with the Spanish registry alone. The Spanish registry therefore appears to support two different values. Since identifying all such cases is too expensive on this instance, we measure them per property by counting the items using the property and the (item, reference) pairs for which one reference supports more than one value:
SELECT (COUNT(*) AS ?n) WHERE {
SELECT ?item ?r WHERE {
?item p:P23 ?st .
?st ps:P23 ?v ; prov:wasDerivedFrom ?r
}
GROUP BY ?item ?r HAVING (COUNT(DISTINCT ?v) > 1)
}
Nine properties were measured: P16 LOA, P17 LBP, P20 GT Tonnage, P21 other tonnage, P22 safety tonnage GTs, P23 main engine power, P24 auxiliar engine power, P32 status on national registry, and P49 IMO. The label lookup failed three times on the Neptune instance, so the labels above were read from the dataset’s own query service, which returns all 60 property labels in a single request even though it cannot answer the reification queries. Labels are quoted as the registry spells them. The result is highly concentrated: P23 has 27,863 conflicting pairs across 173,412 vessels, corresponding to 16.1%, or approximately one vessel in six. Every other property measured sits at 1% or below.

(item, reference) pairs as a share of the items carrying each property, and the same nine properties by absolute pair count on a log scale. P23 main engine power accounts for 96.5% of the 28,877 pairs found; the remaining eight properties hold 1,014 between them. Source: per-property census against the Neptune import, 8 August 2026.This concentration suggests that the issue is specific to engine power rather than a general integration problem. However, the graph does not provide enough information to determine its cause. Because the references contain no retrieval date, three explanations remain possible: an internally inconsistent registry, records retrieved at different times, or an integration step that mapped different source fields to the same property. Distinguishing between these explanations would require additional provenance information, particularly retrieval dates.
The scope of the analysis is also limited: only nine properties were measured. The result should therefore be treated as a targeted finding rather than a graph-wide measure. The concentration in P23, however, provides a specific case for further investigation in a future release.
What the Metrics Do Not Catch
During the RQSS evaluation and subsequent analysis of the EFFKG, we identified several limitations in the dataset:
Outbound interlinking is effectively absent. Running SELECT * WHERE { ?s wdt:P19 ?o } in the query editor over approximately 190,000 entities returns only two rows.

The two rows correspond to Q1 Metre, linked to Wikidata’s Q11573, and Q3 Tonne, linked to Q191118; both are units of measurement. No vessel, port, maritime district, or FAO fishing area contains an outbound link. Moreover, both links use Wikidata /wiki/ page URLs rather than entity URIs, thus, they do not directly connect to Wikidata’s RDF representation. Under the 5-star Linked Data model, the EFFKG therefore does not achieve the fifth-star criterion of linking its data to other datasets.
Placeholder literals pass syntactic validation. For example, Q1000 records its IMO number as "-", which is syntactically valid but semantically empty. The Vessel EntitySchema (E2) defines the expected structure of vessel entities, but this case also requires a constraint on the value itself: an IMO number should be a string conforming to the expected format, rather than an arbitrary string. Without such semantic or pattern constraints, values such as "-" remain valid. Detecting these cases therefore requires domain-specific validation, supporting the use of SHACL validation over the data model alongside the quality assessment.
country of registration is represented as a literal rather than an entity. The value "ESP" is stored as a string, so country-based traversal and entity-level joins cannot be performed directly. This ontology modeling decision therefore prevents an entire class of queries that would otherwise be expressible over the graph.
Duplicate statements survive the import. Q1000 contains two independent statement nodes for has fishing gear category = Q546, with the same value and the same source.
The two deployments have diverged. The public query service exposes 16 reference nodes, including six with no incoming statements, compared with 13 in the imported copy; their source URLs also differ. Consequently, querying the live endpoint and using the DOI-released dump can yield different provenance information. RQSS does not assess this form of divergence, although it is particularly relevant for a dataset intended for citation. We couldn’t find which copy is newer from the available evidence.
A Critical Assessment of RQSS
The evaluation also identified limitations in the current RQSS implementation. In particular, the framework does not consistently distinguish an empty result from a measured value. In statistical terms, an empty set of observations is not equivalent to an observation of zero, and neither should be interpreted as evidence for a positive or negative result.

0,0,1, while Metric 12 reports that ref_sharing.csv was written although the file is absent from the directory.In one experiment, we pointed the extractor to the public endpoint, which does not expose the reification types. The same behavior can also be reproduced with an invalid endpoint, such as by omitting the /sparql suffix. Four related behaviors were observed:
- Metric 32 returns
0,0,1for an empty extraction. The0/0case is resolved as 1, producing a perfect score without observations. - Metric 12 reports successful output although no output file is produced. The process prints
DONE.and exits with status 0, while the expected file is absent. - Metric 14 returns 1.0 when all facts fall into the not-found category. With no observations remaining in the denominator, the metric again produces a perfect score.
- The extractor suppresses query failures. A failed query produces an empty
.datafile and reportsDONE., allowing subsequent metrics to proceed as if the extraction had succeeded.
These behaviors can cause an absence of evidence to be interpreted as positive evidence. RQSS could instead distinguish between successful measurements, valid empty results, and failed extractions, similar to the distinction made in our critical examination of the TEG benchmark, where citation edges that were structurally present but contained no textual evidence were identified explicitly. In particular, undefined cases such as 0/0 should be reported as undefined rather than assigned a numerical score.
A Fair Scorecard
Of the four RQSS metrics completed on the EFFKG import, three provide interpretable results, while Metric 14 primarily exposes a limitation of the current RQSS implementation rather than a property of the dataset.

f8f6c70 against the Neptune import, 9 August 2026.| Metric | Dimension | Score | Interpretation |
|---|---|---|---|
| 1, Availability of External URIs | Availability | 0.769 | Thirteen external URIs were identified; three were unavailable on the measurement date. The same run scored 0.692 one day earlier. |
| 12, Ratio of Reference Sharing | Conciseness | 1.0 | Correctly reflects the high degree of reference sharing, although this pattern is expected in a bulk-integrated graph. |
| 32, Diversity of Reference Properties | Representational consistency | 0.923 | Thirteen reference nodes use a single property, source; no dates are present on the references. |
| 14, Human-Added References | Believability | 1.0 / 0.0 | The result is inconclusive: the score changes with the classification of the importing account, and 134 facts were excluded because their histories could not be retrieved. |
Overall, the EFFKG provides substantial provenance coverage. References resolve to external registries in most cases, and the provenance layer is explicitly represented and queryable. However, the graph generally identifies which registry supports a statement rather than which record and at what time. This limits the verification of individual facts. For example, Part 1 identified a vessel with two engine-power values supported by the same registry; the graph records the registry but provides no retrieval date or record-level identifier with which to resolve the discrepancy.
This distinction is more important than the overall coverage percentage. A provenance layer that identifies a source registry and one that identifies the specific source record and retrieval time may have identical coverage while providing very different levels of verifiability.
The EFFKG nevertheless provides an unusually useful basis for this type of assessment. Enrique Rodríguez-Martín, Jorge Álvarez-Fidalgo, Manuel Luna, and Jose Emilio Labra-Gayo of the WESO Research Group at the Universidad de Oviedo published the data under a DOI together with an open schema, an open integration pipeline, and references attached to the statements. These characteristics make the provenance layer sufficiently explicit to support independent analysis and provide a useful case study for provenance modeling in the Semantic Web community.
The same analysis can be applied to other knowledge graphs. The paper defines the RQSS metrics, the RQSSFramework repository provides their implementation, and the EFFKG Review repository provides the exact commands, queries, and CSVs behind every number in this series. The implementation issues identified above should be considered when interpreting the results. Before running the metrics, it is also useful to inspect the provenance layer directly as RDF using gdotv, the graph data explorer, and the query guardrails where the data is integrated into a pipeline. Direct inspection can reveal structural issues that a numerical quality score may not capture.
References
- Hosseini Beghaeiraveri, S.A., Gray, A. and McNeill, F. RQSS: Referencing quality scoring system for Wikidata. Semantic Web 15(6), 2024, 2419 to 2475. doi:10.3233/SW-243695.
- RQSSFramework, CC0-1.0. Wikibase-generic patches for Metric 14 in commit
f8f6c70. - Zaveri, A. et al. Quality Assessment for Linked Data: A Survey, Semantic Web Journal.
- Rodríguez-Martín, E., Álvarez-Fidalgo, J., Luna, M. and Labra-Gayo, J.E. EFFKG: European Fishing Fleet Knowledge Graph, version 1.0.0, 2026. Zenodo, DOI 10.5281/zenodo.20286320. Licensed CC BY 4.0.
- EFFKG source repository and data model, WESO Research Group, Universidad de Oviedo.
- EFFKG Review queries and scripts, the companion repository for this series. This post’s metric run is
run-all.sh, its standalone queries arequeries/part3-*.rq, and its figure data is underfigures/. - Wikibase RDF dump format, MediaWiki documentation.
- PROV-O: The PROV Ontology, W3C Recommendation.
- Wikimedia Foundation User-Agent Policy.
- 5-star Linked Data, Tim Berners-Lee’s deployment scheme.




