EFFKG Review, Part 2: 4 Queries That Expose the Fishing Fleet Graph

EFFKG Review, Part 2: 4 Queries That Expose the Fishing Fleet Graph, on the gdotv blog

Twenty-six statements describe one Spanish fishing vessel, and every statement has provenance. This is Q1000, the vessel NUEVO SANTA MARIA, returned by a single query without post-processing. Its 26 statements cover 24 distinct properties, with one statement node per assertion, and each statement has at least one reference. All references resolve to two nodes: the European Commission’s Fleet Register and the Spanish national fleet registry.

Trace 26 provenance statements for one vessel back to two registry sources in the EFFKG graph

Entity Q1000 with its 26 statement nodes, the two reference nodes those statements derive from, and the two registry URLs the references resolve to.

This example illustrates an important distinction in the EFFKG: the breadth of the statement layer and the breadth of the provenance layer are different. A graph may contain many statements while relying on a small number of sources to support them.

The European Fishing Fleet Knowledge Graph (EFFKG) is designed around explicit provenance and structured source integration. Part 1 of this EFFKG Review measured the imported copy at 31,698,550 triples covering 190,615 typed entities. This part examines four design features that are particularly relevant to understanding how the graph represents its data:

  1. Statement-level provenance, illustrated by the example above.
  2. Temporal qualification, which represents changes in vessel information over time.
  3. Standards-grounded hierarchy, which preserves the FAO fishing-area hierarchy.
  4. Cross-class connectivity, which determines how these entities can be connected across the graph.

The first three show how the EFFKG’s modeling choices support provenance-aware analysis. The fourth identifies a limitation in the graph’s connectivity and the extent to which further analysis can be performed.

How to Run These Queries

All four queries run against the AWS Neptune import described in Part 1, rather than the public query service, for reasons discussed in the final section. Each query also requires the prefix declarations below. These prefixes must be declared explicitly even when using the dataset’s Query Service, because this Wikibase instance does not automatically resolve its domain-specific prefixes, unlike the Wikidata Query Service.

PREFIX wd:       <https://effkg.wikibase.cloud/entity/>
PREFIX wdt:      <https://effkg.wikibase.cloud/prop/direct/>
PREFIX p:        <https://effkg.wikibase.cloud/prop/>
PREFIX ps:       <https://effkg.wikibase.cloud/prop/statement/>
PREFIX pq:       <https://effkg.wikibase.cloud/prop/qualifier/>
PREFIX pr:       <https://effkg.wikibase.cloud/prop/reference/>
PREFIX prov:     <http://www.w3.org/ns/prov#>
PREFIX rdfs:     <http://www.w3.org/2000/01/rdf-schema#>

Each query below is also published as a standalone .rq file in the EFFKG Review repository, under queries/part2-*.rq, with the prefix block already attached so the file runs as-is against any SPARQL 1.1 endpoint holding the dump.

Provenance Is a Layer, Not a Column

We start with a query with an unbound predicate in the middle of it:

SELECT * WHERE {
  VALUES ?v { wd:Q1000 }
  ?v    rdfs:label           ?vesselName ;
        ?prop                ?stmt .
  ?stmt prov:wasDerivedFrom  ?ref .
  ?ref  pr:P46               ?source .
  FILTER(lang(?vesselName) = "en")
}

The query retrieves each property of the vessel represented by a statement node, then traverses from the statement to its reference and from the reference to its source. This traversal is possible because Wikibase represents each statement as a distinct node. By contrast, the truthy form, wd:Q1000 wdt:P23 176.52, represents the fact as a single edge and carries no provenance information, whereas the reified representation uses four triples to represent the same value while retaining its attribution.

Follow four triples from a vessel item through statement and reference nodes to the source registry URL

The four triples behind one value on Q1000: p:P23 to the statement node, ps:P23 to the literal 176.52, prov:wasDerivedFrom to the reference node, pr:P46 to the source URI. Statement and reference identifiers truncated.

The query returns 37 bindings covering 26 distinct statements across 24 properties. P23 (main engine power) and P25 (fishing gear category) each have two statements. The former represents the conflicting values discussed in Part 1, while the latter is valid because a vessel may use multiple gear categories. Reification alone cannot distinguish these cases; interpreting statement multiplicity therefore requires domain knowledge.

All 26 statements have at least one reference, and 11 have two. Multiple references supporting the same statement provide corroboration, consistent with the provenance model of PROV-O. Across the imported copy, 326,623 of 3,707,591 referenced statements have more than one reference, or approximately one in eleven. The same aggregate query against the public query service returned Unknown Error: ECONNRESET, so this figure is based on the import.

The result also illustrates the concentration of the provenance layer: 26 statements across 24 properties are supported by only two reference nodes, while Part 1 identified 13 reference nodes in the entire graph. The implications of this concentration are examined in Part 3. A graph visualization makes this structure more apparent than a results grid, where the same 37 bindings reveal the counts but not the relationship between statements and sources.

Qualifiers Turn Static Facts into Time Series

The EFFKG represents changes in the number of vessels associated with a port as separate, temporally qualified statements rather than by overwriting the previous value. Santa Pola (Q284) provides a complete example: its P45 value is recorded once for each year from 2006 to 2025, with no gaps.

SELECT * WHERE {
  wd:Q284 rdfs:label ?portName ;
          p:P45      ?stmt .
  ?stmt ps:P45 ?count .
  OPTIONAL { ?stmt pq:P44 ?date }
  OPTIONAL { ?stmt pq:P33 ?since }
  FILTER(lang(?portName) = "en")
}

Read 20 years of dated vessel-count statements fanning out from one EFFKG fishing port node

Port of Santa Pola fanned out into 20 dated vessel-count statements, with the count on the edge and the qualifier date on the target node.

The graph records the series and does not explain it, and neither will this post, since nothing in the EFFKG attributes the 2010 trough to a cause and reading one into it would amount to inventing data. Qualification is also the thinnest of the four design commitments, which is worth stating plainly rather than discovering later in an analysis. Seven properties appear in the qualifier position across the imported copy, for 102,955 usages in total:

Qualifier Label Usages
P13 IRCS 52,865
P33 since 30,731
P55 country of registration 13,989
P44 date 5,180
P47 start date 130
P48 end date 57
P32 status on national registry 3

The most frequently used qualifier in the graph is not temporal. P13, the International Radio Call Sign assigned to a vessel, together with country of registration and status on national registry, accounts for approximately two thirds of all qualifier usages. These properties describe vessel attributes that could instead be represented as statements in their own right, indicating a modeling choice that should be considered when interpreting qualifier-based queries. The four temporal qualifiers account for the remaining 36,098 usages across 3,708,013 statements, meaning that fewer than 1% of statements carry temporal qualification. Temporal coverage is therefore limited, and analyses requiring historical or time-series information should assess qualifier coverage before relying on this representation.

The same breakdown provides a further indication of the limitations of the public query service. It returned a 504 Gateway Time-out, whereas the Neptune import completed the query in approximately one second. Aggregate queries over this dataset are therefore more reliably executed against the import. In such cases, distinguishing a query failure from an empty result is important; query debugging can help establish the cause, while production pipelines should apply appropriate query guardrails.

The FAO Hierarchy Survives the Round Trip

We use a CONSTRUCT query to pull a subtree out of the FAO major fishing area classification:

CONSTRUCT {
  ?child  wdt:P36    ?parent .
  ?child  rdfs:label ?childLabel .
  ?parent rdfs:label ?parentLabel .
}
WHERE {
  ?child (wdt:P36|wdt:P38)+ wd:Q14 .
  ?child (wdt:P36|wdt:P38)  ?parent .
  ?child  rdfs:label ?childLabel .
  ?parent rdfs:label ?parentLabel .
  FILTER(lang(?childLabel)  = "en")
  FILTER(lang(?parentLabel) = "en")
}

The + is a SPARQL 1.1 property path over an alternation, because the EFFKG splits the upward link into two properties, P36 division of and P38 subarea of. The CONSTRUCT template then rewrites both as P36, which is why the render below shows a single predicate rather than two, and it is the reason this query needs CONSTRUCT where the others do not.

Explore the FAO Area 27 fishing area hierarchy rendered four levels deep from EFFKG data

The Area 27 subtree returned by the CONSTRUCT above: 95 items joined by 95 edges, four levels deep at the deepest branch.

Rooted at Q14, Area 27, the result is 95 items and 95 edges reaching four levels below the root:

  • 14 subareas directly under Area 27
  • 38 divisions under those
  • 37 nodes at the next level, 20 of them labeled Subdivision and 17 still labeled Division
  • 4 at the deepest level, including Division 27.3.d.28.1 and Division 27.5.b.1.a

The hierarchy is not uniform in depth or branching. For example, Division 27.5.b.1.a is four hops below the major area, whereas Subdivision 27.9.b.1 is three. Branch sizes also vary substantially: Subarea 27.7 and 27.3 each have 16 descendants, 27.1 and 27.2 have two, while 27.11 and 27.13 have none.

The latter two also expose an important structural property. Both have parent links to Area 27 and Area 34, so the query returns Area 34 even though it is outside the subtree rooted at Area 27. The hierarchy is therefore not a strict tree: two of the fourteen subareas have multiple parents. This distinction matters for downstream traversal and visualization, where a hierarchical layout such as gdotv’s graph data explorer represents the structure more directly than a force-directed layout.

Where the Classes Actually Connect

The next question is how far this connectivity extends across entity classes. The fourth query therefore follows a vessel through its relationships to both ends of the graph:

SELECT * WHERE {
  ?v    wdt:P11 ?port ; wdt:P41 ?area ; rdfs:label ?vl .
  ?port wdt:P3  ?dist ; rdfs:label ?pl .
  ?dist wdt:P3  ?prov ; rdfs:label ?dl .
  ?prov wdt:P3  ?reg  ; rdfs:label ?prl .
  ?reg  rdfs:label ?rl .
  ?area rdfs:label ?al .
  OPTIONAL { ?area wdt:P38 ?sup . ?sup rdfs:label ?sl . FILTER(lang(?sl) = "en") }
  FILTER(lang(?vl)  = "en") FILTER(lang(?pl) = "en") FILTER(lang(?dl) = "en")
  FILTER(lang(?prl) = "en") FILTER(lang(?rl) = "en") FILTER(lang(?al) = "en")
  FILTER(?port = wd:Q284)
}
LIMIT 50

See 50 vessels join the Spanish administrative chain and FAO Area 37 in the EFFKG graph

Fifty Santa Pola vessels, each holding one edge into the Spanish administrative chain and one into FAO Area 37.

All 50 rows have the same structure apart from the vessel. Each vessel links to the Port of Santa Pola through P11 and to Area 37 through P41. From the port, three P3 hops reach the Maritime District of Santa Pola, the Maritime Province of Alicante, and the Autonomous Region of Comunidad Valenciana, using the spelling present in the imported copy noted in Part 1. The OPTIONAL pattern never binds because Area 37 has no supra-area, so the FAO branch terminates after one hop, while the administrative branch extends four hops from the vessel. The longest path therefore contains five edges between Comunidad Valenciana and Area 37, with the vessel providing the only connection between the two hierarchies.

This establishes the structure of the connection and its prevalence requires aggregate counts:

SELECT (COUNT(*) AS ?n) WHERE { ?s wdt:P11 ?o }   # vessel to port
SELECT (COUNT(*) AS ?n) WHERE { ?s wdt:P41 ?o }   # vessel to fishing area
SELECT (COUNT(*) AS ?n) WHERE { ?s wdt:P51 ?o }   # vessel to vessel category
SELECT (COUNT(*) AS ?n) WHERE { ?s wdt:P25 ?o }   # vessel to gear category
SELECT (COUNT(*) AS ?n) WHERE { ?s wdt:P3  ?o }   # administrative, upward
SELECT (COUNT(*) AS ?n) WHERE { ?s wdt:P52 ?o }   # subclass of
Link Property Public endpoint Imported copy
vessel to gear category P25 0 368,502
vessel to vessel category P51 0 187,600
vessel to port P11 0 124,776
vessel to fishing area P41 0 23,830
administrative, upward P3 0 430
subclass of P52 0 13

The zeros in the third column are real! The same count over wdt:P35, the type property used in Part 1 to census the graph, returns only 1,626 entities on the public query service, compared with the 190,591 typed entities reported there. Both measurements were made in August 2026. The cause of this discrepancy cannot be established from the service, so the analysis does not assume one. The practical consequence is that the public truthy layer is not currently reliable for aggregate analysis; the counts in the right-hand column therefore use the import. This is consistent with Part 1’s observation that the endpoint does not expose statement and reference types through the truthy layer.

Part 1 counted 185,047 fishing vessels and 3,147 support vessels in the import, giving 188,194 vessels in total. Against this population, P41 provides 23,830 vessel-to-area links, so at most about one vessel in eight has such a link, and fewer if some vessels belong to multiple areas. The cross-class path described above is therefore not representative of all vessel records. Similarly, P3 provides only 430 links in a graph containing 1,998 ports and 155 administrative units, indicating that most ports do not connect to the administrative hierarchy through this property.

The administrative path also terminates earlier than the hierarchy suggests. The next step is to determine whether the properties that should continue upward point to entities or to literal values:

SELECT ?p (COUNT(*) AS ?total)
       (SUM(IF(isLiteral(?o), 1, 0)) AS ?literals)
       (SUM(IF(isIRI(?o), 1, 0))     AS ?iris)
WHERE { VALUES ?p { wdt:P55 wdt:P56 } ?s ?p ?o . }
GROUP BY ?p

On Neptune, P55 (country of registration) returns 185,707 values, all literals, while P56 (administration responsible for registration) returns 30,529, also all literals. The public endpoint returns no rows for either property, consistent with the results above. Thus, although almost every vessel records its country, none links to a country entity: "ESP" is stored as a string. The administrative chain therefore terminates at the autonomous region, and cross-country traversal cannot be expressed in this model. This ontology modeling decision has a direct analytical consequence: controlled values represented as literals cannot participate in entity-based traversal or joins.

A similar limitation arises for fishing gear categories. The 368,502 gear edges point to only 71 categories, averaging approximately 5,200 edges per category. A sample of 500 vessels illustrates the resulting structure:

CONSTRUCT {
  ?vessel wdt:P25    ?cat .
  ?cat    rdfs:label ?catLabel .
}
WHERE {
  { SELECT DISTINCT ?vessel WHERE { ?vessel wdt:P25 [] } LIMIT 500 }
  ?vessel wdt:P25    ?cat .
  ?cat    rdfs:label ?catLabel .
  FILTER(lang(?catLabel) = "en")
}

Compare fishing gear category hubs absorbing a thousand vessel edges in the EFFKG knowledge graph

A thousand gear-category edges over 511 items. Two categories, Gear not known and Purse Seine, absorb nearly all of them.

Every category is therefore a high-degree hub, making any two vessels that share a category two hops apart. This is structurally correct but provides little additional semantic information: centrality will rank gear categories highest, and community detection will primarily recover the existing ISSCFG classification. The distinction between vocabulary structure and domain structure was also central to our critical examination of the TEG benchmark. The same precaution applies here: category edges should be projected out before applying graph algorithms, or the results should be interpreted explicitly as properties of the vocabulary rather than of the fleet.

What These Queries Do Not Show

The four queries confirm four explicit modeling commitments: statements carry sources, qualifiers represent dates, the FAO hierarchy is preserved, and vessels connect the administrative and FAO hierarchies. These structures are present and measurable, but their presence alone does not establish provenance quality. NUEVO SANTA MARIA has 26 statements supported by two reference nodes, while the entire EFFKG contains only 13 distinct reference nodes. These references identify the source registry but provide no retrieval date, record identifier, or version. Thus, complete statement-level coverage does not necessarily provide sufficient information to verify individual facts.

Part 3 examines this distinction directly by applying the published RQSS referencing-quality framework to the same imported copy and reporting its metrics as defined by the framework.

The connectivity query also provides a simple way to examine other parts of the graph. Replacing wd:Q284 with another port and running the query against an import of the EFFKG distribution shows how far that entity connects into the surrounding hierarchies. gdotv can connect to the same RDF graph through its generic RDF connector and provide a visual representation of the resulting structure.

References

Try gdotv with Amazon Neptune (RDF)

Query, visualize and model your Amazon Neptune (RDF) data on your laptop, in your own cloud or across the whole team.

gdotv Desktop

The ultimate graph database IDE.

Install and connect your first graph database in 2 minutes.

gdotv Team

The unified workbench for graph data teams.

Deploy to your team in under 15 minutes. Marketplace or self-hosted.

gdotv Enterprise

The universal graph intelligence platform.

For organization-wide deployments. Bring graph data directly to your analysts.