EFFKG Review, Part 1: How to Query the Fishing Fleet Knowledge Graph

EFFKG Review, Part 1: How to Query the Fishing Fleet Knowledge Graph, on the gdotv blog

The European Fishing Fleet Knowledge Graph (EFFKG) was released in May 2026 by Enrique Rodríguez-Martín, Jorge Álvarez-Fidalgo, Manuel Luna and Jose Emilio Labra-Gayo of the WESO Research Group at the Universidad de Oviedo. It integrates maritime, fisheries and administrative data drawn from international standards, European registries and Spanish national sources into a single Wikibase instance. We flagged its release in the Weekly Edge at the time and invited readers to do something with it, and this three-part EFFKG Review is our own exploration of the dataset. Every query and script behind the series is published in its companion repository, so each number below can be reproduced rather than taken on trust.

Two access routes exist and this post takes both: the public SPARQL endpoint, which requires no setup, and a local import of the published RDF distribution into AWS Neptune. Running an unfiltered query against the imported copy shows the structure of the data:

Explore the first imported EFFKG batch rendered as a graph in gdotv against AWS Neptune

An unconstrained triple pattern over the first Turtle batch of the distribution, loaded into AWS Neptune and rendered in gdotv. Statement nodes outnumber Items roughly three to one, and wasDerivedFrom is among the most frequent predicates.

The dominant nodes visible in this view are the statement nodes, listed as Statement in the legend, which form the Wikibase reification layer. They are the reason the dataset can attribute nearly every fact it asserts to a source. Reification is explained in the next section.

Running the same queries against the two routes produces materially different answers, which is a finding in its own right and one that constrains every measurement in the remainder of the series.

What the European Fishing Fleet Knowledge Graph Contains

The dataset is dominated by vessels. Grouping entities by their instance of (P35) statements produces the following census, computed separately for the public endpoint and the imported copy:

Class Public endpoint Imported copy
Fishing Vessel 185,025 185,047
Vessel supporting fishing related activities 3,145 3,147
Fishing Port 1,997 1,997
Fishing Area 114 114
Maritime District 108 108
Fishing Gear Category 71 71
Fishing Vessel Category 48 48
Maritime Province 30 30
Vessel supporting fishing related activities Category 18 18
Spanish Autonomous Community 17 17
Material 9 9
On board Equipment 4 4
Vessel Category 2 2
Category 2 2
Industrial Port 1 1
Total typed entities 190,591 190,615

“Vessel” classes account for more than 98% of the typed entities in both copies. The remaining classes are smaller but provide much of the graph’s structural organization: ports, the Spanish maritime administrative hierarchy, FAO fishing areas, and the classification vocabularies derived from the FAO’s ISSCFV and ISSCFG standards for vessel and gear types.

The census also reveals two differences between the copies. The imported copy contains 22 additional Fishing Vessel entities and 2 additional support vessels, for 24 more typed entities in total. Its class labels also differ: it uses Maritimal District and Maritimal Province, whereas the public endpoint uses Maritime District and Maritime Province. This indicates that there have likely been editorial changes in the live instance since version 1.0.0, including the spelling correction and the removal of a small number of vessel records. This does not establish which copy is more accurate; it shows that the published distribution is a snapshot of an instance that has subsequently changed.

For version 1.0.0, the authors report 190,628 entities, 3,707,656 statements, and 4,033,845 provenance references. Neither copy reproduces the reported entity count exactly, and the gap is a difference in what is being counted rather than missing data. The census above counts entities that carry an instance of statement, and the class names in its first column are themselves Items in the graph. Most of those class Items carry no instance of statement of their own, so the census does not count them even though the authors’ entity total does.

The Wikibase Statement Model

The EFFKG is built on Wikibase, the platform that also underlies Wikidata. Wikibase and Wikidata are distinct: the former is the software platform while the latter is a public instance of that platform, and the EFFKG is a separate Wikibase instance. They nevertheless share an RDF representation in which a statement is represented as a node rather than as a single triple, a modeling approach commonly referred to as reification. A direct claim in this representation is expressed as a single edge:

wd:Q1000  wdt:P23  176.52

This statement links entity Q1000, the Spanish fishing vessel NUEVO SANTA MARIA, to property P23 (“main engine power”). You can view all the EFFKG properties here.

That form, which the Wikibase documentation calls the truthy claim, discards everything about where the value came from. The reified form routes the same fact through intermediate nodes instead:

Trace how the Wikibase RDF data model links an item to statement, value and reference nodes

The Wikibase RDF mapping, with relevant edges shown in red. p: reaches the statement node, ps: its simple value, prov:wasDerivedFrom the reference node, and pr: the source that reference names. Source: Michael F. Schönitzer, Wikimedia Commons, CC BY 4.0.

The four red edges in this representation support the provenance queries used throughout this series. For Q1000 and its engine power property P23, p:P23 connects the item to the statement node, ps:P23 connects the statement to the value 176.52, prov:wasDerivedFrom connects the statement to a reference node, and pr:P46 connects the reference to its source. The wd:, wds:, and wdref: prefixes shown in the diagram are those used by Wikidata; the EFFKG uses the same structure with its own domain. Resolving these prefixes against the wrong domain is the first issue addressed by Route One below.

The remaining elements of the model are used less directly in this series. Qualifiers attach to statement nodes through pq:, which provides the temporal information examined in Part 2. Quantities and dates that require structured representation use value nodes reached through psv:, including the TimeValue node shown in the schema derived from the imported copy. The prov:wasDerivedFrom predicate is defined by PROV-O, providing a standard provenance vocabulary rather than a Wikibase-specific predicate.

On one hand, this representation increases the number of triples required for each attributed statement, but on the other it makes provenance explicitly addressable and allows conflicting values to be represented without resolving them during ingestion.

The benefit of this representation is apparent at the level of an individual vessel. The NUEVO SANTA MARIA records two values for main engine power: 176.52 and 240.0. Neither value is discarded, and neither is marked as preferred or deprecated under the Wikibase ranking mechanism. Both are retained as independent statements, each with an attribution to its source registry.

Compare two conflicting engine power values for one fishing vessel in the EFFKG provenance graph

Two statements, two values, three provenance edges: the 176.52 statement (lower pink node) is supported by both registries, while the 240.0 statement (upper pink node) is supported by the Spanish registry alone.

In the gdotv visualization, each pink edge is a prov:wasDerivedFrom link from a statement node (pink) to a reference node (green), and each dark green edge carries that reference on to the registry URL it names (purple).

The structure of this disagreement is more informative than the disagreement itself, because the reified representation distinguishes three cases.

The first case is corroboration. The statement with value 176.52 has two reference nodes: one resolves to the European Commission’s Fleet Register and the other to the Spanish national fleet registry. Two independent sources supporting the same value is a primary purpose of provenance-aware modeling.

The second case is explicit disagreement, which the model supports but this vessel does not show. If each registry reported a different value and both were retained with their own attributions, that would also be a valid representation, and it would preserve source-level disagreement that selecting a single value would discard.

The third case is what the second statement actually shows. The value 240.0 has a single reference, again pointing to the Spanish registry, so the same source appears to support both 176.52 and 240.0 for the same property of the same vessel. That is distinct from both corroboration and disagreement between sources: it may indicate an inconsistency in the source registry, or an integration issue involving different source fields or retrievals.

The graph cannot distinguish between these explanations because each reference records only the source registry, with no retrieval date, record identifier, or file version. The usefulness of reification therefore depends on the information represented by the reference node itself. Part 3 examines how widespread this limitation is and what it implies for referencing quality across the dataset.

The reification layer also affects how the graph is represented by graph visualization tools. Applying gdotv’s data model viewer to the imported RDF and deriving its schema produces the Wikibase meta-model rather than the domain model of the fisheries data:

Inspect the Wikibase data model gdotv discovers from the imported EFFKG RDF batch files

The schema discovered from the first imported batch: Item, Statement, Reference, TimeValue and schema:Dataset. Derived from one batch only, so it is indicative rather than complete.

Vessel, Port, and Fishing Area do not appear in the derived schema because they are not RDF/RDFS/OWL classes. They are Items linked through P35 (“instance of”), a property local to the EFFKG instance rather than a term from RDF, RDFS or OWL, so their types are represented in the instance data rather than in a schema that discovery can read from rdf:type and rdfs:subClassOf. Schema discovery therefore exposes the Wikibase reification structure instead of the fisheries domain classes. This distinction between schema-level types and types encoded in instance data is a recurring issue in RDF modeling, discussed in ontology modeling for RDF and property graphs.

The unconstrained graph render shows the same effect at the instance level: statement nodes greatly outnumber Items, reflecting the reified representation required for provenance. The pattern is also visible in the TEG benchmark series, where the structure observed after import can differ substantially from the conceptual model described by the dataset. Deriving the schema in two steps from the loaded data therefore provides a useful way to establish this difference before querying the graph.

Route One: The Public Query Service

The EFFKG provides a query service at https://effkg.wikibase.cloud/query, but two implementation details are important when accessing it programmatically. First, this URL serves the browser interface rather than the SPARQL API. A SPARQL client must therefore use https://effkg.wikibase.cloud/query/sparql. The distinction is not obvious from the project documentation or the Zenodo record, both of which present the shorter URL as the SPARQL endpoint.

Second, unlike Wikidata’s query service, this instance does not implicitly declare the standard prefixes used in Wikibase queries. Using prefixes that resolve to Wikidata can therefore produce empty results rather than an explicit error. Queries against the EFFKG should consequently declare their prefixes explicitly against https://effkg.wikibase.cloud/:

PREFIX wd:   <https://effkg.wikibase.cloud/entity/>
PREFIX wdt:  <https://effkg.wikibase.cloud/prop/direct/>
PREFIX p:    <https://effkg.wikibase.cloud/prop/>
PREFIX ps:   <https://effkg.wikibase.cloud/prop/statement/>
PREFIX pr:   <https://effkg.wikibase.cloud/prop/reference/>
PREFIX prov: <http://www.w3.org/ns/prov#>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

The Query Service also supports multiple views, including geospatial queries comparable to those available through Wikidata’s Query Service. Ports record coordinates through P42, allowing the service’s geospatial functions and map rendering directive to be used directly:

Map fishing ports within 150 km of Aviles using the EFFKG Query Service geospatial view

Ports within 150 km of the Port of Avilés, rendered through the query service’s built in map view. Source: EFFKG Query Service, CC BY 4.0.

For work outside the browser, gdotv connects to the endpoint through the RDF (Other) connection type, which supports any SPARQL 1.1 endpoint and exposes it in the query editor alongside the graph view. This generic support across the triplestore range allows the EFFKG to be accessed without a vendor-specific connector:

Configure the gdotv RDF connection to reach the public EFFKG SPARQL endpoint with /sparql appended

Connecting to the EFFKG endpoint in gdotv using the generic RDF connector, with /sparql appended to the documented URL.

Route Two: Importing the Distribution

The published distribution is a 1.4 GB archive on Zenodo, released under CC BY 4.0. After extraction, it contains 3,813 Turtle files rather than a single RDF file. This structure is suitable for bulk loading but less convenient for browser-based upload. In our setup, staging the files in S3 took longer than the Neptune load itself, and individual batches occasionally failed because of transient network errors and required retries:

Retry one failed EFFKG Turtle batch upload into the Amazon S3 staging bucket for Neptune

Re-uploading a single Turtle batch to the S3 staging prefix after a network failure. The batch shown is 1.4 MB.

The mechanics of Neptune bulk loading are covered in detail in Part 2 of the TEG benchmark series and in our Neptune setup walkthrough, so they are not repeated here. The relevant parameter is format=turtle, and the loader is pointed at the S3 prefix holding all 3,813 batches. When I conducted this import, the load completed without error:

Read the AWS Neptune bulk loader status reporting a completed EFFKG knowledge graph load

Final loader status: 3,813 feeds completed, 143,921,443 records processed, no parsing, datatype or insert errors.
Loader field Value
feedCount (LOAD_COMPLETED) 3,813
totalRecords 143,921,443
totalDuplicates 112,982,109
totalTimeSpent 2,042
parsingErrors 0
datatypeMismatchErrors 0
insertErrors 0

Two figures from the load report are noteworthy. The load completed in 2,042 seconds, approximately 34 minutes, which provides a useful baseline for loading a distribution of this size. More notably, the loader reports 112,982,109 duplicate records out of 143,921,443 processed, meaning that approximately 78% of processed records duplicated statements already encountered in the distribution. This may reflect repeated vocabulary and label triples across batches rather than redundancy in the underlying graph; the loader output does not distinguish between these cases. The absence of parsing, datatype, and insertion errors across all 3,813 files nevertheless provides evidence that the published export is syntactically consistent.

What the Import Reveals

Running the same three counting queries against both copies produces the following:

Measure Public endpoint Imported copy
Total triples 26,090,047 31,698,550
Nodes typed wikibase:Statement 0 3,708,013
Nodes typed wikibase:Reference 0 13
Statements with prov:wasDerivedFrom 3,707,230 3,707,591

The zeroes in the second and third rows are the key result. The public query service returns no nodes for ?s a wikibase:Statement or ?s a wikibase:Reference, whereas the imported copy returns 3,708,013 and 13 respectively. The published RDF therefore contains type assertions for statement and reference nodes that are not exposed by the served graph.

This is documented platform behavior rather than an EFFKG publishing decision. The Wikibase RDF dump format specification lists it under WDQS data differences: types for wikibase:Item, wikibase:Statement and wikibase:Reference “are currently omitted for performance reasons”. The filtering happens in the munger, the preprocessing step that rewrites a dump before it is loaded into the query service backend, and it has been the default since 2015, when T115242 added an opt-out flag precisely because “for certain use cases, like object counting and comparison, it is desirable to retain these”. The same specification recommends matching statements with wikibase:rank and references with prov:wasDerivedFrom instead of rdf:type.

The practical consequence is important. Queries that depend on these type assertions return empty results from the public endpoint without raising an error, while the same queries work against the imported RDF. This affects analyses based on the Wikibase RDF model, including the RQSS framework examined in Part 3. The endpoint remains usable through property paths, as shown by the 3,707,230 statements returned in the final row, but the silent absence of type assertions makes this limitation difficult to detect. The 5,608,503-triple difference between the two copies is consistent with the missing type assertions and other administrative triples not exposed by the Query Service.

The imported copy also establishes two important properties of the provenance layer. Of 3,708,013 statement nodes, 3,707,591 carry a reference and 422 do not, giving a coverage of 99.9886%, consistent with the 99.989% reported by the authors. Provenance coverage is therefore genuinely close to complete.

However, those 3,707,591 referenced statements resolve to only 13 distinct reference nodes. Wikibase content-hashes reference nodes, so identical reference content is represented by a single shared node, and the provenance layer therefore identifies only thirteen distinct reference contents across the whole dataset. Whether this concentration is a limitation or an appropriate representation of a bulk integration is the subject of Part 3, where the same imported copy is evaluated using a published referencing-quality framework.

Which Route for Which Question

What you want to do Route
Query the graph as currently published and maintained Public endpoint
Use geospatial rendering without additional tooling Public query service
Work against a pinned, citable snapshot of version 1.0.0 Imported copy
Address statement or reference nodes by type Imported copy
Explore visually and produce figures Either, through gdotv

The two copies serve different purposes. The endpoint represents the current state of the live instance and may include changes made after the published release. The distribution is the citable snapshot and preserves the complete RDF model, including the type assertions required to query the reification layer directly. Analyses of statement and reference nodes should therefore use the distribution, whereas analyses of the current fleet register should use the endpoint. Published measurements should state which copy they use, since the two already differ in basic counts such as the number of typed entities.

Returning to NUEVO SANTA MARIA and the disagreement between its engine power values, both copies preserve the conflict and attribute both values to the Spanish registry, so the conflict is not an artifact of versioning. Neither records when the underlying data was retrieved.

Part 2 of this series examines the structures the graph does represent, including provenance, source reconciliation, temporal qualification, and connectivity between vessels, ports, and fishing areas.

Part 3 then evaluates the reference layer using a published referencing-quality framework.

The endpoint is publicly accessible without credentials. To reproduce the queries, use any SPARQL 1.1 client or gdotv with https://effkg.wikibase.cloud/query/sparql and run the provenance query for NUEVO SANTA MARIA. The counting queries above, and every other query in the EFFKG Review, are in the series repository.

References

Try gdotv with Amazon Neptune (RDF)

Query, visualize and model your Amazon Neptune (RDF) data on your laptop, in your own cloud or across the whole team.

gdotv Desktop

The ultimate graph database IDE.

Install and connect your first graph database in 2 minutes.

gdotv Team

The unified workbench for graph data teams.

Deploy to your team in under 15 minutes. Marketplace or self-hosted.

gdotv Enterprise

The universal graph intelligence platform.

For organization-wide deployments. Bring graph data directly to your analysts.