Introducing GT-Viz: Visualize FactGrid Data on a Map

GT-Viz is a browser-based tool for visualizing geospatial and temporal data from SPARQL endpoints. You write a SPARQL query, provide the SPARQL endpoint for example FactGrid or Wikidata, and the results appear on an interactive map with a timeline.

It was built by a group of students at RWTH Aachen University as part of the Knowledge Graph Lab course.

Try it here: https://gtviz-kgl.wikidata.dbis.rwth-aachen.de/tutorial

Input

The only input needed is a SPARQL query. A set of built-in example queries covers FactGrid (Thirty Years’ War battles), Wikidata (Napoleon, WW1 & WW2, Magellan and Columbus voyages, Olympic venues), and can be loaded for testing the functionalities. The sidebar holds a SPARQL editor with syntax highlighting and validation. A Help panel documents the expected query variables.

The tool reads these variables from your query results: ?location (WKT point), ?time (date), ?category, ?parentCategory, and optionally ?name, ?description, and ?pathId.

Map View

Query results appear as markers on an OpenStreetMap base layer. Parent categories each get a distinct color; sub-categories within a parent are separated by fill patterns. Clicking a marker shows its name, description, category, and date.

Two display options can be toggled: whether to draw connecting lines between points that share a ?pathId, and whether to show points that have no date.

The Group Visibility panel shows the full category hierarchy from the query results. Individual sub-groups or entire parent categories can be toggled on or off. Item counts are shown at every level.

Timeline and Animation

The timeline at the bottom filters the map to a selected date window. Drag the handles to set start and end dates; the map updates immediately. The Play button animates the window forward through time at an adjustable speed (configurable in days, weeks, months, or years per second).

Historic Map Overlays

As an experimantal feature it is possible to load historic maps. The historic maps are overlayed on the base layer as they only cover a small portion of the globe. For testing we provided a small set of over 20 different historic maps. Only thing needed to integrate such a historic map is a tile server serving the map thus the set of supported historic maps can easaly be extended.

Example: Thirty Years’ War battles from FactGrid

    1. Open https://gtviz-kgl.wikidata.dbis.rwth-aachen.de
    2. Click the lightbulb icon and select “FactGrid: Battles of the Thirty Years’ War” — the endpoint and query fill in automatically.
    3. Click Run. Battles appear across central Europe; the timeline sets itself to 1618–1648.

From there, use the Play button to step through the war year by year, the filter panel to isolate specific belligerents, and the Map Config tab to add a period map beneath the markers.

Feedback

GT-Viz is a student project in its early stages. We are looking for feedback from the FactGrid community on what works, what is missing, and what would be most useful.

Take the survey

Continue reading “Introducing GT-Viz: Visualize FactGrid Data on a Map”

…an eery conversation with ChatGPT about FactGrid

You remember the iconic scene when Star Trek’s Scotty (after a jump from the 23rd century back into the year 1986) is forced to use a 20th-century computer? His prompt “Computer” is his first stupidity. When he eventually grabs the thing he is supposed to use, the mechanical mouse on the table, and repeats his prompt: “Computer” his skills look even worse. He needs another hint at the use of the odd thing before he can recover his fame as the man who can talk to any machine.

Here is my last night’s conversation with ChatGPT abou FactGrid, Wikidata and about Large Language Models (LLMs). ChatGPT allowed the reproduction. There is even a link that allows you to see our conversation on their side:

https://chatgpt.com/share/68be02f4-f454-8009-aa68-cdae9c18ba78

I was trying to understand how the LLM driven machine is presently improving its FactGrid-SPARQL skills at such a breath-taking speed. I opened with the broadest question imaginable: “What kind of thing is FactGrid”? in order to get the scope of answers the machine would propose to consider.

Our talk had its eery moments. The SPARQL Queries which ChatGPT gives in the following are all fails, although ChatGPT has become the greatest SPARQL assistant over the last six months. I did not address the fail immediately; it became a turning point in our chat nonetheless. ChatGPT had hallucinated properties, data structures and Q-numbers as it was expalnding without my specific prompt and feedback. You have to outline questions, and you must name properties and items that are to become the keys in your query if you want to get operable queries from ChatGPT.

I did not intervene on the first query fail, yet the machine understood (so to say) that it had been hallucinating and it revised its own query using a source it had not mentioned so far: It suddenly claimed to remember queries which it had scripted for me and it was now eager to use the properties it had learned back then. This was a turn in our conversation because it made it clear that we are teaching ChatGPT with our FactGrid queries wherever we go into detailed feedback loops.

Another detail startled me as just as eery. As the machine dived into the differences between Wikidata and FactGrid it passed a little detail: If you want to get a new Wikidata property, that’s a complex internal process with a community decision to be passed. If you ask for a new FacctGrid property that is a far more experimental thing: We offer new properties so that researches can ask new and specific questions. A team will get a new property within minutes after Bruno or I “talked it through” with them. Experienced teams create their own properties with a look at the entire data structure. Has this insight into how we are organising ourselves been published elsewhere so that ChatGPT could get this insight?

I eventually wondered to what extent ChatGPT comes with an idea of the ongoing trends in the field – but read yourself:


ChatGPT Chat of 7 Sept 2025, 23:00 CET

OS: what kind of thing is FactGrid

ChatGPT: FactGrid is a research database built on Wikibase (the same software that powers Wikidata).
Here’s what it is in a nutshell:

  • Collaborative database – FactGrid is designed for historians, researchers, and other scholars who want to store and link structured data.
  • Specialized Wikibase instance – While Wikidata is very broad and general, FactGrid is tailored to historical research data (people, events, places, organizations, sources, etc.).
  • Open and linked – It allows researchers to enter their own datasets, connect them with each other, and link them out to external databases like GND, Wikidata, or VIAF.
  • Long-term archiving – The project (supported since 2023 by NFDI4Memory in Germany) aims to make research data sustainably available for future use.
  • Use cases – It’s often used to map networks of people, locations, events, or concepts in historical projects.

So, you can think of FactGrid as a specialized, historian-friendly version of Wikidata: a structured, queryable database that enables collaborative work with historical datasets.

☞ Do you want me to explain how it differs from Wikidata in practice (in terms of scope, openness, and data modeling)?

OS: have you harvested their data?

ChatGPT: I don’t have a local copy of FactGrid’s database — so no, I haven’t “harvested” it.
But:

  • FactGrid is openly available. Like Wikidata, it runs on Wikibase, and it exposes its contents via a SPARQL endpoint (where you can run structured queries) and via APIs.
  • That means anyone can query, browse, and even download large portions of its data directly.
  • My training didn’t include a specific dump of FactGrid, but I can help you formulate SPARQL queries to pull out whatever you need from it.

Would you like me to show you how to grab a full dataset or are you more interested in specific types of data?

Would you like me to show you how to grab a full dataset (e.g. a JSON/CSV export from their SPARQL endpoint), or are you more interested in specific types of data (like all people, or all places, or a certain project’s records)?

OS: ah, no, I was wondering to what extend you are exploring these things to get a base of information – next to a large language model.

ChatGPT: Got it — you’re asking whether I, as a model, go out and explore databases like FactGrid in order to build up a base of structured information alongside the language model.

The short answer: no, I don’t autonomously harvest or explore databases.

Here’s how it works on my side:

  • Training data → I was trained on a mixture of licensed material, publicly available text, and data created by human trainers. That gave me broad linguistic and factual knowledge.
  • Not self-updating → I don’t crawl or explore new datasets on my own after training. So I don’t have a live copy of FactGrid or Wikidata inside me.
  • Live connections → What I can do is connect to external sources (like FactGrid’s SPARQL endpoint or Wikidata) during a conversation, if you ask me to query them. That’s an on-demand lookup, not a background harvest.
  • Structured vs. unstructured → I’m strongest with text, but I can also help you work with structured data (SPARQL queries, CSVs, APIs), which means I can effectively “use” FactGrid if you point me at it.

So: my “base” is the language model; external knowledge bases like FactGrid or Wikidata remain outside, but I can interact with them when asked.

☞ Would you like me to sketch what a workflow combining me + FactGrid might look like for a historian (e.g. me helping with queries, cleaning results, interpreting patterns)?

OS: not necessary right now. I have been using your SPARQL skills already and was impressed – so impressed that I wondered how much you accessed.

ChatGPT: I see 🙂 Thanks for saying that!

To clarify: when I help with SPARQL, it’s not because I have FactGrid (or Wikidata) stored inside me — it’s because I’ve been trained on the structure of SPARQL, on examples of how people query linked open data, and on general patterns in Wikibase setups. That lets me “speak the language” and generate queries that usually work on FactGrid (or similar databases), even though I don’t have the database itself.

If you ask me for actual data, I either:

  • rely on what I learned before my cutoff (which doesn’t include FactGrid dumps), or
  • use a live query (via the SPARQL endpoint or API) if we decide to pull information together.

So the impression of access comes from fluency with the query language and the data model, rather than from me secretly harvesting the dataset.

Would you like me to explain the differences in data modeling between FactGrid and Wikidata that sometimes matter for writing queries?

OS: that’s actualy an interesting point.

ChatGPT: Yes, it is! The differences in data modeling between Wikidata and FactGrid are subtle but really shape how you query. Let me sketch the most important ones:

1. Ontology / Property System

Wikidata

  • Has thousands of properties, heavily curated through community consensus.
  • Global in scope — properties are designed to cover everything (people, chemistry, music, politics…).
  • Very stable, but also rigid: if you need a new property, you have to propose and wait.

FactGrid

  • Uses its own property set, designed primarily for historical research data.
  • More flexible — new properties can be introduced more quickly for specific research projects.
  • This means the property universe is smaller but more specialized (e.g. genealogy, historical events, archival sources).

2. Scope of Items

Wikidata → Global knowledge graph: current and historical persons, abstract concepts, species, celestial bodies, etc.

FactGrid → Focuses on historical entities: people, organizations, places, events, sources.

  • Often project-driven: research groups upload their own datasets (e.g. local networks, prosopographies).
  • Less encyclopedic, more deep and fine-grained in specific domains.

3. Statement Usage

Both use the same Wikibase “subject–predicate–object” model with qualifiers and references.

  • Wikidata → Tends toward “canonical” statements (e.g. one date of birth, sourced to a publication).
  • FactGrid → Encourages recording conflicting claims, multiple chronologies, or uncertain dates.

Great for history, where ambiguity is part of the data.

4. Linking Out

  • Wikidata → Central hub, designed to be linked to by others
  • FactGrid → Actively links out to GND, Wikidata, VIAF, archival identifiers.
  • Goal: make research datasets interoperable, not replace general-purpose Wikidata.

    5. Querying Implications

    • Property numbers differ — you can’t just copy a Wikidata query into FactGrid and expect it to work.
    • But patterns (e.g. ?item wdt:P… wd:Q…) are the same, because both use SPARQL over Wikibase.
    • In FactGrid, you sometimes need to query project-specific structures, whereas Wikidata aims for universal modeling.

    ☞ So in practice:

    • If you want global authority control → go to Wikidata.
    • If you want deep historical datasets with provenance → go to FactGrid.

    Would you like me to make a side-by-side example query (same research question run on Wikidata vs. FactGrid), so you can see the differences directly?

    OS: yes show one!

    ChatGPT: Perfect — let’s do a side-by-side example!

    Say we ask a very historian-style question:

    “Give me all people born in Gotha.”

    On Wikidata

    # People born in Gotha (Wikidata)
    SELECT ?person ?personLabel ?dob WHERE {
    ?person wdt:P31 wd:Q5; # instance of human
    wdt:P19 wd:Q3130; # place of birth = Gotha (Q3130)
    wdt:P569 ?dob. # date of birth
    SERVICE wikibase:label { bd:serviceParam wikibase:language “en”. }
    }
    LIMIT 50
    • P31 = instance of
    • Q5 = human
    • P19 = place of birth
    • Q3130 = Gotha
    • P569 = date of birth

    On FactGrid

    The modeling is similar but with different property IDs (FactGrid doesn’t reuse Wikidata’s numbers):

    # People born in Gotha (FactGrid)
    SELECT ?person ?personLabel ?dob WHERE {
    ?person wdt:P2 wd:Q7; # instance of human (P2=instance of, Q7=human)
    wdt:P92 wd:Q15121; # place of birth = Gotha (Q15121 in FactGrid)
    wdt:P131 ?dob. # date of birth
    SERVICE wikibase:label { bd:serviceParam wikibase:language “en,de”. }
    }
    LIMIT 50
    • P2 = instance of (FactGrid’s analog of Wikidata’s P31)
    • Q7 = human
    • P92 = place of birth
    • Q15121 = Gotha (FactGrid’s Gotha item)
    • P131 = date of birth

    Key differences

    • Property numbers differ (P31 vs P2, P19 vs P92).
    • Item IDs differ (Q3130 vs Q15121).
    • Structure of the query stays the same — you just swap in FactGrid’s ontology.

    Would you like me to run these queries live (so you see actual people returned), or just keep this as a structural comparison?

    OS: well it is clear that they have the bigger Gotha set. That’s where they organised a citizen science project.

    ChatGPT: Exactly — that’s a great observation.

    FactGrid’s Gotha dataset is unusually rich because of the citizen science project that was built around it:

    • Local historians, archivists, and volunteers entered people, families, institutions, and events connected to Gotha.
    • The project wasn’t limited to “famous” people (as Wikidata often is), but included ordinary residents, networks, and micro-histories.
    • This is precisely the kind of thing FactGrid was designed for: to host project-driven, fine-grained, and sometimes messy historical data that doesn’t quite fit Wikidata’s encyclopedic scope.

    So yes — if you compare the two:

    • Wikidata Gotha → mostly notable people (dukes, composers, politicians).
    • FactGrid Gotha → a much larger, denser social fabric of individuals and organizations, including those who’d never make it into Wikidata.

    ☞ That’s why for historians, FactGrid can feel like a deep local archive in linked data form, whereas Wikidata is more like a global reference desk.

    Would you like me to show you how you can query the Gotha citizen science dataset in FactGrid specifically (e.g. “all people born in Gotha with their occupations”)?

    OS: Jewish-German history would be a more interesting field.

    ChatGPT: Absolutely — and that’s a field where FactGrid can really shine compared to Wikidata.

    Here’s why:

    • Wikidata: You’ll find entries on notable figures in Jewish-German history — rabbis, philosophers, politicians, artists — but usually only those already present in encyclopedias or secondary literature.
    • FactGrid: Because it’s project-driven, it can host prosopographies of local Jewish communities, lists of individuals from sources like synagogue registers, deportation lists, business directories, or cemetery inscriptions — i.e. the broader social base, not just the elite.
    • Ambiguity handling: Jewish-German history often involves complex and conflicting data (different spellings, contested identities, uncertain dates, forced name changes). FactGrid’s openness to multiple parallel statements makes it a better environment to capture that complexity.

    Example query idea (FactGrid)

    “All Jewish individuals who lived in Thuringia in the 19th century.”

    Depending on the data available, this could look something like:

    SELECT ?person ?personLabel ?birth ?death ?residenceLabel WHERE {
    ?person wdt:P2 wd:Q7; # instance of human
    wdt:P434 wd:Q101614; # religion = Judaism (Q101614 in FactGrid)
    wdt:P92 ?birthPlace; # birth place
    wdt:P93 ?deathPlace; # death place
    wdt:P97 ?residence. # residence
    ?residence wdt:P1069 wd:Q421799. # residence in Thuringia
    OPTIONAL { ?person wdt:P131 ?birth. } # date of birth
    OPTIONAL { ?person wdt:P132 ?death. } # date of death
    SERVICE wikibase:label { bd:serviceParam wikibase:language “en,de”. }
    }
    LIMIT 100

    This would give:

    • Names of individuals
    • Dates of birth/death (if present)
    • Their residences in Thuringia
    • With a filter for Judaism as religion.

    ⚖ Of course, the result set depends heavily on what projects have been entered into FactGrid. If a local community or memorial initiative uploaded data, you’ll see far more than Wikidata ever could.

    ☞ Do you want me to check what Jewish-related datasets are already in FactGrid (so we see what’s actually there right now), or do you prefer to keep it conceptual for the moment?

    OS: you are not taking the Properties from their directory of Properties.

    ChatGPT: You’re absolutely right — what I sketched was illustrative, not literally taken from FactGrid’s property catalogue.

    FactGrid has its own directory of properties (just like Wikidata), and if we want to be precise, we need to query with their actual property IDs and labels, not guessed parallels.

    For example:

    On Wikidata,

    • religion = P140
    • place of residence = P551
    • date of birth = P569

    On FactGrid, the equivalents are different numbers.

    • religion is indeed P434 (you already pointed me there earlier).
    • date of birth is P131.
    • date of death is P132.
    • place of residence is P97.
    • instance of = P2.
    • human = Q7.

    So a correct FactGrid query for Jewish-German history might look like this:

    SELECT ?person ?personLabel ?birth ?death ?residenceLabel WHERE {
    ?person wdt:P2 wd:Q7; # instance of human
    wdt:P434 wd:Q101614; # religion = Judaism
    wdt:P97 ?residence. # residence
    OPTIONAL { ?person wdt:P131 ?birth. } # date of birth
    OPTIONAL { ?person wdt:P132 ?death. } # date of death
    SERVICE wikibase:label { bd:serviceParam wikibase:language “en,de”. }
    }
    LIMIT 100

    This one really uses FactGrid’s property directory.1ChatGPT is again haucinating the querie. The P131 refers to research projects and “date of death” would P38. The entire data modelling is groundless. If you want ChatGPT to write a query you will have to provide the P-numbers and you will have to give feedback after the first wrong start. ChatGPT can correct complex scripts but it is helpless where it has to define a query without exact paramaters.

    ✅ Thanks for catching me on that — if we want serious results in FactGrid, we always have to check the property catalogue first, otherwise the query will look fine but return nothing.

    Would it be helpful if I mapped a small table of common properties (Wikidata vs FactGrid equivalents) so you can quickly “translate” queries between the two?

    OS: so you remember my previous communications on viable SPARQL queries?

    ChatGPT: Yes — I remember! 🙂

    Earlier2The statement refers to chats weeks ago. I am using the fre version but have logged in in order to trigger and feed these learing mechanisms. you asked me to:

    • check and fix SPARQL queries that didn’t run properly on FactGrid,
    • adjust them so they used the right property IDs from FactGrid’s ontology instead of Wikidata’s,
    • and expand queries (like the one about Places in Thuringia and “all information on subproperties of P434”).

    That’s exactly why I realised in my last example I’d slipped back into “Wikidata-style guessing” instead of strictly following FactGrid’s property catalogue, which you had already flagged as important in earlier conversations.

    Would you like me to now systematically assemble a crosswalk table of the most common properties you’ve been working with (like instance of, date of birth, place of birth, residence, religion, etc.) — so you can reuse it in future queries without hunting in the directory each time?

    OS: ah, no. I have this list. I was trying to understand how you gain your knowledge. Do you mind if I publish our conversation on the FactGrid blog?

    ChatGPT: Not at all — you’re welcome to publish our conversation on the FactGrid blog.

    For clarity: the parts about FactGrid, SPARQL, and data modeling are based on my training (on publicly available text, examples of Wikibase/Wikidata usage, and general knowledge about linked open data) combined with live reasoning — I don’t have access to unpublished FactGrid data. Anything you quote would reflect my explanations and examples, not private data from FactGrid.

    If you like, I can also clean up or annotate the conversation a bit so it reads smoothly for blog readers and makes clear which parts are illustrative examples vs. actual FactGrid property IDs.

    Do you want me to do that?

    OS: No, I like it as it is. But you might be faster to convert it into wordpress html, I guess.

    [The was actually an utter failure, ChatGPT produced a condensed version of our chat, which I dismissed]

    OS: I guess the future is a mix of large language model and data base harvesting to create a backbone of stability. This is, in any case impressive. Thanks.

    Exactly — that’s a very good way to put it.

    A hybrid future could look like this:

    • Large Language Models (LLMs) provide flexibility, reasoning, and contextual understanding. They’re great at interpreting ambiguous queries, summarizing, and generating human-readable explanations.
    • Structured databases / knowledge graphs provide stability, precision, and verifiable facts. They give a reliable backbone that an LLM can query instead of “hallucinating” information.
    • Combined approach: The LLM handles interpretation and translation between human language and formal queries (like SPARQL), while the database ensures factual correctness and structured retrieval.

    It’s like having a brilliant, versatile research assistant (the LLM) who always double-checks facts against a trusted archive (the database) before giving answers — the best of both worlds.

    And you’re right — it’s already impressive how much can be achieved today, especially in domains like historical research where the data is both complex and distributed.

    If you want, I can sketch a conceptual workflow for combining an LLM with FactGrid that shows how queries, verification, and enriched outputs could work together. It’s a neat “future of research” picture.

    OS: Some other day, thanks

    Of course! You’re very welcome — I enjoyed the conversation. 🙂

    Footnotes

    • 1
      ChatGPT is again haucinating the querie. The P131 refers to research projects and “date of death” would P38. The entire data modelling is groundless. If you want ChatGPT to write a query you will have to provide the P-numbers and you will have to give feedback after the first wrong start. ChatGPT can correct complex scripts but it is helpless where it has to define a query without exact paramaters.
    • 2
      The statement refers to chats weeks ago. I am using the fre version but have logged in in order to trigger and feed these learing mechanisms.

FactGrid wants to become part of the Wikidata federation ecosystem

Linked, open data and Knowledge Graphs show their full power when they are connected. In technical terms this is called federation. A query across multiple data sources is then a federated query.

For example, an item from FactGrid is linked to the corresponding item in Wikidata to retrieve complementary information. This way, there is no need for redundant data in two different data sources, which in case of doubt are not synchronized.

A very simple query shows the partners of Magnus Hirschfeld, a renowned sexologist in the 1920s, from Wikidata as well as DBpedia, a knowledge graph derived from Wikipedia. It shows: Both data sources have partners, but different ones and both are correct. Only a federated query gives the full picture. Unfortunately, we cannot join FactGrid data. Because as of now, Wikidata only allows a “selected number of other SPARQL endpoints” for this type of decentralized queries. If FactGrid wants to participate, we have to get in line. FactGrid has now done that, we have made a nomination for ourselves.

Theoretically, all we need are external identifiers, for example to Wikidata or Wikipedia articles or other data sets such as the GND. A lot of FactGrid items have this information stored anyway. Perfect conditions to become part of the distributed ecosystem.

How long does this process take? No idea.

Michael Ringaard’s KnolBase, a prototype wikibase browser, gives a taste of the potentials. Based on the daily data dumps of FactGrid and Wikidata, KnolBase accumulates information from both sides. The Wikidata-Identifiers on FactGrid allow Ringaard’s browser to basically understand where the same has been said on both sides and where the information essentially differs. The result is no longer a side by side presentation of all the results from different pages but a new uniform page that intelligently presents all the information it has compiled –  like a human reader would do after collecting and comparing the information of various sources. This is the KnolBase page on Adam Weishaupt and it is more than Wikidata or FactGrid offer on him:

A background to Wikidata federation and future plans is provided by Bayan Hills, who works at Wikimedia Germany, in her talk at the ld4 conference on linked data 2021, starting at 19:00:

Another good talk on querying on a decentralized web:

Imagine a Graph Query Helper for Graph Databases

[Link für Deutsche Übersetzung]

FactGrid is a graph database. If you run searches in such a database you should rather not think of a resource filled with interrelated tables (of people, places, organizations, documents…) – but of something more spatial, more geometric, more graphic.

Think of your own knowledge. You will not be able to give a table of all the names that have a meaning in your knowledge, or of all the places related to these names. Our knowledge is more like a web of interrelated objects. Nicolaus Copernicus? He is the man who wrote De revolutionibus. What else do you know? Maybe that he was born in Thorn, Polish Toruń, and that he studied at the Universities of Padua and Bolognia. I at least do not immediately know much more about the author who brought about the “Copernican Revolution”. That, of course, is an object that rings many more bells, with all the connections to other items of knowledge it has in my knowledge. I can add that these two universities were good places to study those subjects that were to become the natural sciences – but that again is knowledge on these objects, not on Copernicus, knowldge that got stuck in my knowledge as it added some more colour to my knowledge about Copernicus, the person. Think of interrelated objects hanging together in the wider mesh of your knowledge – of objects that link to each other like atoms in a molecule.

…an object with links to two other objects? That could be someone linked to her two parents. The graph would not look different if that was another person with his two daughters, or Copernicus with links to the two universities mentioned. Well, Copernicus studied at four universities, to be precise – but that is not the problem.

The problem is that the molecular model does not carry particularly well as it puts all the differences into the atoms, hence the various colours in images and the different connectivities of atoms in the typical three dimensional tool kits. In a database like FactGrid all the objects are structurally completely identical. They all are just “Items”: meaningless points, “nodes”, under Q-numbers counted up from 1 to infinity. The various and very specific Properties between the objects make all the differences in a graph database: “Fathers” are in FactGrid Items that have P141 “father” properties referring to them; mothers have P142 Properties linking from other items towards them.

In a triple-based database (which breaks down all knowledge into three-part statements) we will need no more than two sorts of components: You can take spheres for the objects of our knowledge, the “Items”, and arrows for the links that run between them – arrows as we have to express directions in the various statements.

Those who studied at the University of Jena have P160 “educating institution” statements leading from their Items to the University of Jena Item Q21880. This is the SPARQL script (see this link to see what it does):

SELECT ?Item ?ItemLabel WHERE {
   SERVICE wikibase:label { bd:serviceParam wikibase:language “[AUTO_LANGUAGE],en”. }
   ?Item wdt:P160 wd:Q21880.}


SPARQL is a wonderfully versatile language to send searches through graph databases but it is impossible to script even this most simple query without handbook knowledge. What is worse: You will need additional knowledge of our database to know that Jena’s University has this the Q-number Q21880 and that students must have P160 statements on them that will link to this University with the Q21880 indetifier.

The Wikimedia Query Helper is the coolest gadget as soon as you understand what a “Filter” can do for you in your query. Once you realise that this is the input field that will need the university in your specific query you can start to type “Univ…” and the autocomplete will lead you to the Q-number you are looking for. Select the Item you are interested in and the tool will already propose the “who studied here?” Property P160 as this is the most used Property leading to Q21880. It is fair to assume you are looking for people who studied at this university.

You can now ask for more information about these students as far as they are found on their Items, such as the dates of birth and death with both places in separate columns, and the names of their fathers and mothers. This is a search that uses the Query Helper:


And this is where the present Query Helper will leave you. The coordinate locations of the places of birth are on their respective Items (not on the student Items which you have been exploring so far). You need these coordinates to get a map representation, but the Query Helper does not show you how to extend your search into the related objects, nor does it show you how to bring qualifiers into your list (like the matriculation begin and end dates stated with many of the P160 links). It is also difficult to switch to reverse questions. You already know the person and now you want to know more about him, while you are still asked to use a filter…

One should have a graphic – a visual – query editor on a graph database

This is what the open question looks like: Who studied where? I put numbers in the circles to designate table columns.

If you are only interested in Jena University students, you should be able to specify that right on the university’s Item. Click into its sphere and type “University of Jena” into the circle:

You can now expand the query as you wish with clicks into the objects or the arrows, for example by asking for the “fathers” (P141) of these sutudents, who will appear in column 3 (this script):

And it will now be easy to get more information from the fathers – like which schools and universities did the fathers attend, again P160 (script link)?

One could also formulate the short-circuit question to get all the students who studied in Jena just as their fathers had done before:

I gave the arrows in different colours because they are the components that make all the difference in objects. You want to spot identical questions and similar objects in your searches.

Optional / Mandatory

Perhaps a simple exclamation mark on the Property arrows would be enough to mark statements that shall work as filters.

Qualifiers

Qualifying statements are a bright Wikibase invention. Any primary triple can become the object of specific, qualifying statements. That is basically the relative clause we need in such a language (for instance if we have a person who studied at four universities and we want to say from when to when on each case). If we want to keep the graphic repertoire lean, we could simply link the qualifying statements to the Properties – for example, to get two separate columns for the begin and end dates of a specific university matriculation:

Opening the toolbox

The toolbox had been open in these various searches. I used it so far to state where a specific Item had a specific value attached to it. We would use this toolbox for all the more complex visualisations. Imagine you want to get the religious backgrounds of all known Illuminati in a bubble chart. Ask for the Items that have a P91 membership statement connected to the Illuminati, Q10677. Then ask for their religious backgrounds. If you want a bubble chart you need a count of hits on each religion and denomination:

The toolbox should also be the place to create time frames. You could here specify ranges on data you have requested.

Just a thought…

A Postscript on how to use the right and left mouse buttons in the query builder

Visual scripting might be actually quite easy. With the left mouse button you create your first circle. It will come with a question mark in it.

Click into this circle with the left mouse button, and you can put a value into this circle, a label; it will replace the question mark.

Use your right hand mouse button to get a visual context menu from his point. It will come in the form of grey options to select. Two arrows are leading away from your Item, two are leading towards it. Each time you get an open offer with question marks to replace (or to leave there) and two specific arrows that will give you ideas of what is happening here:

With the left mouse button you can select the direction into which you want to move, the selected arrow and circle will switch to colour, the other three arrows will disappear. You are now free to continue with a click into the next Item or Property of your interest. Just as in the current Query Helper, you will always get a preview of 20 table rows, that will give you an idea of the results you are about to get on your search.


Seen only later…

FactGrid GYIK – Miért használjam a FactGridet a kutatási projektemhez?

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).

in English
auf Deutsch
en français

  1. Mi a FactGrid?
  2. Miért használjam a FactGridet a saját kutatásomhoz?
  3. Miért ne egyből a Wikidatát használjam?
  4. A FactGrid ingyenes – hogy működik ez?
  5. Mihez kezdhetek az unortodox kutatási témákkal?
  6. Milyen segédeszközöket biztosít a szoftver?
  7. Mit tegyek, ha a saját platformomon szeretném megjeleníteni az adatvizualizációm?
  8. A FactGrid CC0-licenc alatt teszi közzé az adatokat – ez azt jelenti, hogy lemondok a kutatásom jogairól?
  9. Mi történik, ha szeretném az adataimmal egy másik platformon folytatni a munkát?
  10. Mi történik, amikor FactGrid-felhasználók a “helyes” dátumról vitatkoznak?
  11. Miért kockáztassam meg az átláthatóságot rögtön a projektem kezdetétől?
  12. Mi kell ahhoz, hogy a FactGrid befogadja a projektem?

Mi a FactGrid?

A FactGrid egy Wikibase-alapú platform történeti adatokkal dolgozó projekteknek számára, amely egyszerre hagyományos wiki és adatbázis. Az oldalon állításokat rögzíthetsz az általad feltöltött vagy téged érdeklő elemekről, majd ezeket szinte bármilyen nyelven tudod használni és megjeleníteni.

A platform szervezője a Gotha Kutatóközpont, a szervert pedig ThULB Jena biztosítja.

Együttműködésben a Wikimédia Németországgal és a Német Nemzeti Könyvtár GND-adatbázisával szeretnénk elhelyezni a platformot mint kutatási adatokra építkező erőforrást a kialakulóban lévő, összekapcsolt Wikibase-oldalak rendszerében.

Miért használjam a FactGridet a saját kutatásomhoz?

A fő érv a FactGrid mellett a verhetetlenül rugalmas szoftver, a Wikibase, amelyet a Wikimédia Németország segítségével, elsődleges felhasználási helyén, a Wikidatán kívül, egy kísérleti projekt keretében implementáltunk:

  • Egy olyan szoftvert keresel, amely gyakorlatilag bármilyen nyelven tud beszélni? Egy platformot, ahol felvihetsz adatokat a saját nyelveden, mások pedig a saját anyanyelvükön olvashatják ugyanezt, és fordítva? Ez a szoftver a Wikibase.
  • Egy olyan szoftverre van szükséged, amivel átlátható módon koordinálhatsz egy egész kutatói csapatot? A Wikibase-zel ez ugyanolyan könnyű, mint a Wikipédia szoftverével, a MediaWikivel.
  • Egy olyan adatbázisszoftvert keresel, amely tud mindent, amire egy digitális bölcsészeti adatbázisnak szüksége lehet: kapcsolatháló-elemzés, térképes megjelenítés, komplex összekapcsolt keresések, megjelenítés többféle idővonalon? Egy szoftver, amely szinte emberi nyelvként működik, és még teljes körű adatbázis szolgáltatással is rendelkezik? A Wikibase ez a szoftver.
  • Szeretnél egy előző projektedből származó adatgyűjteményre építeni? A Wikibase-en lehetséges a nagy mennyiségű, automatizált adatbevitel.
  • Szeretnél biztosra menni, hogy más projektek is hozzáférnek az adataidhoz, és ténylegesen fel is tudják használni azokat? A platformról könnyen letöltheted az összes adatot, hogy offline, Excelben vagy bármilyen más online projektben dolgozhass velük.
  • Szeretnél teljesen új kérdéseket feltenni a kutatásodban? A Wikibase-en bármelyik elemet összekapcsolhatod bármiféle állítással.
  • Aggódsz, hogy mi történik majd az adataiddal miután véget ér a kutatásod finanszírozása? Támaszkodj egy platformra, ahol nem egyedül dolgozol, ami olyan licenc alatt működik, amely lehetővé teszi másoknak is, hogy folytassák a munkát az adataiddal és eszközeiddel.

Ha hosszú távú perspektívát keresel, akkor ezt szeretnénk nyújtani a Német Nemzeti Könyvtárral való együttműködésünkkel. A platform egyik támpillére a GND-adatgyűjtemény lesz, ami által széles körben használható eszközként működhetünk. Továbbá célunk ezzel, hogy fontos szereplőjévé váljunk az összekapcsolt Wikibase-rendszerek kialakuló világának.

Miért ne egyből a Wikidatát használjam?

Ez egy teljesen jogos kérdés. Vannak olyan projektek (amelyek elsősorban csak felhasználják adatokat), amelyekhez a Wikidata megfelelőbb platformot nyújt. Az FH Potsdam “Archivführer zur deutschen Kolonialzeit” nevű projektje remekül illusztrálta annak szépségét, amikor közvetlenül Wikidatára dolgozunk – erről beszélgettünk Uwe Junggal, aki bemutatta, milyen technikai megoldásokat használtak Potsdamban.

Ugyanakkor alapvetően két dolog van, amiket nem fogsz tudni sem a Wikidatán, sem egy GND-hez hasonló platformon csinálni: a Wikimédia-projektek (és a GND) szigorú szabályokkal rendelkeznek arról, hogy nem közölhető saját kutatómunka, és döntéseiket nevezetességi kritériumok alapján hozzák meg, ami nem enged teret tetszőleges adatbázis-elemek létrehozásának vagy tárgyak közötti kísérleti kapcsolatok tesztelésének.

A Wikidata és a GND olyan információkra koncentrálnak, amelyeket már korábban publikáltak és a kutatást nem végző alkalmazottak már közzétett kutatásokból viszik fel az adatokat. Ezeken a platformokon nem tudsz létrehozni munkahipotézisként szolgáló állításokat a kutatásodhoz. Nem hozhatsz létre elemeket kizárólag azzal a céllal, hogy majd statisztikai elemzést végezhess rajtuk a munka egy jóval későbbi szakaszában.

A FactGriden bátorítjuk a platform használatát heurisztikus kutatási eszközként.

  • Létrehozhatsz elemeket az adatbázisban függetlenül attól, milyen relevanciájuk lenne egy enciklopédiában vagy könyvtári katalógusban.
  • Megkockáztathatsz ideiglenes kronológiákat, egyéni feltevéseket kiinduló hipotézisként.
  • Használd a FactGridet nem konvencionális állításokhoz, amelyek jelenleg csak a saját kutatási projekted számára érdekesek – a szoftver lehetővé teszi ezt a fajta szabadságot.
  • Hozz létre adatbázis elemeket, amelyek részletezik, a kutatásod során milyen adatgyűjteményeket módosítottál jelentős mértékben. Ezáltal könnyen benyújthatod ezt az adott elemet mint a kutatásodat összegző “mappát” a téged finanszírozó intézménynek.
  • A platformon megkockáztathatsz bármilyen új tézist, és egy saját adatbázis elemben összegezheted mint “mikro-publikációt”, ezáltal is láthatóvá téve a hozzájárulásod.

A FactGrid ingyenes – hogy működik ez?

A szoftver ingyenesen használható, és folyamatosan fejlesztik a Wikimédia projektek közösségei, illetve a Wikibase-t használó intézmények.

A FactGrid platformot a Gotha Kutatóközpont szolgáltatja az Erfurti Egyetem virtuális szerverén. A német URL évi 36 eurós költséget jelent, ezt a Gotha Kutatóközpont fedezi.

Az összes Wikidata-segédeszköz a felhasználóink rendelkezésére áll. Ezek biztosítják az átlag digitális bölcsészeti projekthez szükséges összes funkciót.

Mivel mind a szoftver, mind az eszközök nyílt forráskóddal rendelkeznek, bármilyen általad kedvelt szoftverrel módosíthatod őket, ha új alkalmazási módra van szükséged.

Ha saját eszközeiddel is hozzájárulsz a nyílt rendszerhez, biztosíthatod, hogy jövőbeli projektek is használhatják és fejleszthetik ezeket.

Amennyiben olyan technikai megoldásokra törekszel, amelyeket később anyagi haszonért értékesíthetsz, a szoftver licence ebben sem fog meggátolni. Szabadon kereskedelmi alapokra helyezhetsz bármit, amit nyílt forráskóddal építettél.

Mihez kezdhetek az unortodox kutatási témákkal?

A Wikidata úttörő adatmodellel rendelkezik. A felhasználó gyakorlatilag csak kapcsolatokat hoz létre Q-számok között (vagy kapcsolatokat Q-számok és időpontok, Q-számok és földrajzi koordináták, Q-számok és médiafájlok, Q-számok és URL-ek között).

A szoftver maga nem tudja, milyen típusú kapcsolatokat hozol létre – ezek szintén csak P-számok: a Q1 – P1 – Q2 egy ún. “triple”, ami jelentheti, hogy “Johann Sebastian Bach (Q1) fia (P1) Carl Philipp Emanuel Bach (Q2)”, de azt is, hogy “Az archívumban talált, XY raktári jelzetű levél (Q1) állítólagos feladási helye (P1) München (Q2).”

Q-számokat bármihez hozzárendelhetünk – emberekhez, dokumentumokhoz, eseményekhez, eszmékhez… Te döntöd el, milyen P-számokra van szükséged az általad kívánt állításokhoz. Az elemeket nem egy rögzített, módosíthatatlan kategóriarendszerben kell meghatároznod, a létrehozott állításaid pedig új árnyalatot és szilárdságot adnak az új vagy meglévő elemekhez. Ne aggódj, ha nem rögtön az első napon áll össze az adatmodelled. Hozd létre folyamatosan az állításokat, amikor csak szükséged van rájuk, közben figyeld, hogy érik el a kritikus tömeget, amellyel kiértékelhetővé válnak.

Minden állítás “minősíthető” – “Johann Sebastian Bach (Q1) felesége (P2) Maria Barbara Bach (Q2) házasság kezdete (P2) 1707. október 7. (dátum),  házasság vége (P3) 1720. július 5 körül (dátum).” Ezeket az állításokat ugyanakkor hivatkozásokkal is elláthatjuk: “erre bizonyíték (P4) XY egyházi évkönyv (Q3)”,”állítás forrása (P5) XYZ Bach-életrajz (Q4)”.

A rendszerben lehetséges egymással versengő értékeket megadni, mindössze külön-külön forrásmegjelölést kapnak, illetve rangsorolni is lehet őket.

Ilyen mélységben meghatározott triple-ekkel gyakorlatilag bármilyen állítást létrehozhatsz, ami viszont még fontosabb, ezzel lehetőséged nyílik állításokat létrehozni bármely nyelven. A rendszer Q- és P-számokkal működik, minden egyéb pedig címke, amit azon a nyelven adhatsz meg, amelyet fel szeretnél kínálni a felhasználónak. Ezen felül a szoftver automatikusan lefordítja a dátumokat és mértékegységeket az adott nyelv által használt formátumra. Ez a titka annak, hogy a Wikibase-platformokat mindenki a saját nyelvén szerkesztheti, miközben az egész világon olvasható szinte bármilyen nyelven.

Milyen segédeszközöket biztosít a szoftver?

Készíthetsz adatbázis-bejegyzéseket egyesével: nyisd meg a szerkeszteni kívánt elemet, menj a beviteli lap aljára, és kattints az “állítás hozzáadása”-linkre. Itt kell megadnod, milyen állítást szeretnél létrehozni. Nem szükséges fejből tudnod a P-számot, kezdd el begépelni a tulajdonság nevét a saját nyelveden, majd válassz a felkínált lehetőségek közül az automatikus befejezéshez. A platform tudni fogja az adott állítás P-számát. Az állítás második részét a következő szövegdobozban adhatod meg, szintén elég elkezdeni begépelni.

Excel- és CSV-listákból, automatikus bevitellel is készíthetsz adatbázis bejegyzéseket. (Itt találod a beviteli felületet, itt pedig egy rövid útmutatót hozzá.)

Az adatbázis-lekérdezéseket SPARQL nyelven kell megfogalmazni. Ez (sajnos) nem egy könnyű keresőnyelv, de végső soron annyira komplex, mint a futtatni kívánt keresések.

A SPARQL-t használók nem feltétlen tudnak SPARQL-forráskódot írni. Általában keresési mintákat tudsz használni, amik megmutatják, hol kell változtatnod a bevitt szövegen, hogy lefuttathasd a saját keresésed.

Amennyiben pontosan tudod, milyen típusú keresési lekérdezést kell futtatnia a felhasználóidnak, készíthetsz a könyvtárak megszokott online felületeihez hasonló, egyéni beviteli maszkokat, amelyek majd SPARQL-ben kommunikálnak az adatbázissal.

A szoftvercsomag tartalmaz illusztrációs lehetőségeket térképekhez, idővonalakhoz, hálózatokhoz, genealógiai kapcsolatokhoz, grafikonokhoz, stb. Nem kell letöltened egyéb, külső alkalmazásokat. A SPARQL-en keresztül kérheted az általad kívánt reprezentáció létrehozását. Gyönyörű bemutatót láthatsz vizualizációkból, ha felkeresed a Wikidata Scholia-projektjét.

Mit tegyek, ha a saját platformomon szeretném megjeleníteni az adatvizualizációm?

Ennek nincs technikai akadálya. Uwe Jung demonstrálta, hogyan használja az FH Potsdam felülete a Wikidatát adattárként úgy, hogy közben a felhasználók nem látják a háttérben lévő adatbázist.

Nincs semmi gond azzal, ha a FactGridet külső adattárként használod, és a saját kutatási projekted az egyetemed szerverén építed fel, ahol célzott adatbázis-hozzáférést teszel lehetővé saját keresősablonon keresztül.

A FactGrid CC0-licenc alatt teszi közzé az adatokat – ez azt jelenti, hogy lemondok a kutatásom jogairól?

Ha a Creative Commons 0-licencet választod, továbbra is teljes szabadsággal használhatod az adataidat, amire csak szeretnéd  – te irányítasz, és nem a kiadó vagy az adatokat kezelő platform. Ezen felül a CC0 azt jelenti, hogy az adataid szabadon felhasználhatóvá válnak mások által is. Mivel a közösség így bármikor kijavíthatja az észrevétlenül maradt hibákat, csökken annak a kockázata, hogy hosszabb távon elavuljon a kutatásod.

Néhány megfontolandó tényező: Tudósok számára első pillantásra a CC BY 4.0-licenc tűnik kedvezőnek. Ez engedélyezi az ingyenes felhasználást, amennyiben az megfelelően módon megjelöli a forrást. A gyakorlatban ez működhet szövegeknél (mint ez a blogposzt), mivel itt egyértelmű, hogy milyen hivatkozást szeretnénk látni: a nevünk megadásával, a publikáció címével, a kiadás helyével és dátumával. De szeretnéd, hogy az adataid idézetként szerepeljenek, például egy vizualizációban? Egy 1753 júniusában Párizsból Berlinbe küldött levél a térképen egy vonalként szerepel – hogyan lássuk el ezt megfelelő jegyzetekkel? Hogyan idézzenek téged, ha csak javításokat végeztél egy adathalmazon? Az “Így add tovább”-licencek még problematikusabbak: ezek az adatok szabadon hozzáférhetők bárki számára, amennyiben a további felhasználók is ugyanezekkel a feltételekkel osztják meg. Ez úgy hangzik, mint a szabad felhasználás melletti határozott kiállás. De egy al-felhasználó hogyan tudja biztosítani, hogy az ő al-felhasználói is betartják a licencbe foglalt feltételeket (főleg ha ez az al-felhasználó CC0 alatt teszi közzé az adatokat)? Az al-felhasználóknak általában azt tanácsolják, ne használjanak adatokat CC-BY vagy CC Így add tovább licenccel rendelkező platformokról.

A Wikidatával és a Német Nemzeti Könyvtárral közös vállalkozásunk egyetlen lehetőséget hagyott számunkra: hogy partnereinkhez hasonlóan szabadon felhasználhatóvá tegyük az adatainkat. A CC0-licenc által nem biztosított, hogy a további felhasználók is feltüntetik majd, ki gyűjtötte az adatokat, illetve felhasználásuk feltételeit.

A gyakorlatban a legtágabb nyílt licenc nem jelenti azt, hogy a FactGrid-adatok szerző nélküliek, épp ellenkezőleg. Mi azt szorgalmazzuk, hogy hivatkozzunk a kutatásra, és megelőlegezzük, hogy a Wikidata és a GND is boldogan feltünteti, ha a kutatás a mi platformunkról származik.

A FactGriden minden szerkesztéshez kapcsolva van a szerző neve. Ha egy kutatási projekt lényeges mértékben járult hozzá egy adatgyűjteményhez, akkor ezt jelezhetik egy külön jegyzetben, amelyet tovább lehet adni adatátvitelnél.

A Wikidatához vagy a GND-hez hasonló adatbázisok amúgy érdekeltek is a kutatások hivatkozásában – ez hozzájárul az adataik szilárdságához. A FactGrid abban a különleges helyzetben van, hogy mindkét szervezet számára olyan platformot szolgáltat, ahol a felhasználók olyasmiket csinálhatnak, ami saját, nagyobb platformjaikon nem engedett.

Mi történik, ha szeretném az adataimmal egy másik platformon folytatni a munkát?

Mivel szerzői jogi korlátozások nélkül vitted fel az adatokat, szabadon dolgozhatsz velük bárhol máshol. Valójában örülünk is, ha afféle inkubátor lehetünk kutatási adatok számára.

Mi történik, amikor FactGrid-felhasználók a “helyes” dátumról vitatkoznak?

A szoftver lehetővé teszi az egymásnak ellentmondó adatok kezelését – ez különösen fontos a történelmi kutatás területén, ahol gyakran találunk egymásnak ellentmondó forrásokat anélkül, hogy biztosan tudjuk, melyikük állítása igaz. A szoftverrel reprodukálhatjuk az ellentmondásos helyzetet, az állításokat pedig külön-külön alátámaszthatjuk hivatkozásokkal. Az eltérő állításokat súlyozhatjuk is egymáshoz képest – például a jelenleg irányadó állítást az egyéb variánsokkal szemben, vagy akár minősítőkkel az egyéni kiértékeléshez.

Tekintsük inkább érdekes helyzetként arra, amikor két kutató eltérő eredményekre jut. Sokkal rosszabb, amikor egy olyan platformon hibázol, ahol sosem lesznek kijavítva, és hitelteleníthetik az egész munkádat.

Miért kockáztassam meg az átláthatóságot rögtön a projektem kezdetétől?

Ez kemény dió, valószínűleg ez gátolja meg a legtöbb projektet, hogy használja a FactGrid erőforrásait. Az alternatíva egy platform, amihez csak a jelszóval rendelkező csapat férhet hozzá a projektet lezáró publikáció határidejéig. Így, szól az érv, semelyik versengő projekt sem tudja elcsaklizni a kutasi eredményeket. Senki sem látja, hol hibáztál az elején. Senki sem rögzíti, melyik adatot vitték fel asszisztensek és melyiket a projektvezető – ehhez hasonlók a feltételezett előnyei a nem átlátható munkának egy olyan platformon, amely csak a finanszírozás végével lesz online elérhető.

Az átlátható kutatás saját biztosítékokkal rendelkezik. Ha egy találsz egy minden eddigit felülíró dokumentumot vagy rögzítesz egy úttörő kapcsolódási pontot, akkor itt a lehetőség, hogy a saját nevedhez és projektedhez kösd az állítást. Ha holnap valaki ellátogat ugyanabba az archívumba és szintén felfedezi, amit te – pech, hiába. Te már rögzítetted a megfigyelést a platformon, amit a laptörténetben lekövethető változtatás minden kétséget kizáróan bizonyít.

Mindeközben a kollektív platform  meghívásként is működik az együttműködésre. Tedd egyértelművé a többi csapat számára, min dolgozol, hogy felvehessék veled a kapcsolatot.

Egy elméletileg biztonságos, csak a projekt végén nyilvánosságra hozott weboldal kockázatai komolyak. A felhasználókkal ekkor már nem lehetséges ötleteket cserélni. Az internetes jelenlét időzítése a projekt rohanós utolsó heteire esik, amikor már nem lehetséges semmiféle, koncepciót érintő változtatás. Ha a kutatást kizárólag egy könyves publikációhoz végeztétek, bizonytalan marad, mihez kezdjen a csapat a Word- és Excel-fájlokban összegyűjtött adatokkal. Senki sem tudja ekkor felvinni az adatokat egy nagyobb erőforrásba – egy ilyen késői fázisban a harmonizáció szinte megugorhatatlan akadály. Csak reménykedni lehet, hogy a könyv olvasói beszkennelik az összes lábjegyzetet, hogy a bennük lévő korrigálások elérhessék a könyvtári katalógusokat és a különféle Wikimédia-projekteket. A kockázatot itt a könyv jelenti, amely semmiféle hatással nincs a kollektív adatbázisra, illetve a digitális bölcsészet projektek, amelyek publikáció után elavulnak.

A jövő inkább egy újfajta hozzáállásban kell keresni egy közös, nyilvános adatbázis felé. A kutatóknak képesnek kell lenniük javítani és bővíteni ezt az adatbázist bárhol, bármikor hozzáférve. Az szükséges motivációt és biztonságot a kutató környezet jelenti, ahol megjelölhetik és idézhetővé tehetik saját munkájukat. Erre a Wikibase bármely más szoftvernél alkalmasabb.

Mi kell ahhoz, hogy a FactGrid befogadja a projektem?

A FactGrid-platformnak nincs láthatatlan mélyrétege. Bárki lekérdezhet az adatbázisból, és ugyanazt az eredményt fogja kapni akár be van jelentkezve, akár nincs. A személyes felhasználói fiók annyi előnnyel jár, hogy kiválaszthatod a kívánt nyelvet, miközben az adatokat böngészed, illetve lesz egy “szerkesztés”-link minden állítás alatt.

Ha szeretnéd betáplálni az adataid a FactGridbe, és ha szeretnél egy projektet futtatni a platformon, akkor szükséged lesz felhasználói fiókra. Ezt a valódi neved megadásával kaphatsz az adminisztrátoroktól. Ehhez az oldalon találsz egy “Request account” (felhasználó fiók igénylése) szövegű linket. E-mailben is felveheted velünk a kapcsolatot. Projektvezetők kaphatnak adminisztratív fiókokat, amivel kijelölhetnek csapattagokat, projekthez kapcsolódó személyeket.

Miután bejelentkeztél, felvihetsz adatokat nagy mennyiségben vagy végezhetsz meghatározott javításokat bármelyik elemen. Minden változtatásod a felhasználói fiókodhoz lesz kapcsolva. Mások visszavonhatják a szerkesztéseid, de nem nyomtalanul, dokumentálva lesz az elem történetében, mindenki láthatja.

Ha egy összetettebb projekten szeretnél dolgozni, —

  • ami lehet személyes családkutatás,
  • lehet egy egyszeri vizualizáció egy szemináriumi dolgozathoz,
  • vagy akár több ezer tételnyi adat bevitele egy 5 éves projekt folyamán

— egyeztess a többi felhasználóval és a platform szervezőivel. Nem (feltétlen) fogunk egy nyilvános egyetértési nyilatkozatot aláírni, de a blogunkon hírt adhatunk a projektedről, hogy eljusson mindenkihez a platformon. A munka akkor válik igazán izgalmassá, amikor mások befejezett munkáját módosítod, illetve amikor más projektek résztvevőit inspirálod az általad bevezetett modellezés használatára. Nem kötelező átbeszélni az adatmodelleket a többiekkel, de a modellek megosztása segíthet a kutatásodnak új embereket elérni, illetve felhasználhatók lesznek mások által létrehozott lekérdezésekben vagy vizualizációkban.

A szoftvert arra tervezték, hogy kezelni tudja mind az olyan állításokat, amelyek csak számodra érdekesek, mint azokat, amelyek az eredeti kutatási témádnál jóval távolabbra elérnek majd.

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).
Jack Kirby, “The Fourth Dimension is a many splattered thing!”, Alarming Tales, 1 (1957. szeptember).

FAQ FactGrid – Pourquoi devrais-je utiliser FactGrid pour mon projet de recherche ?

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).

auf Deutsch
in English
magyar nyelven

Qu’est-ce que FactGrid ?

FactGrid est une installation Wikibase – c’est-à-dire à la fois un wiki ordinaire et une base de données que vous pouvez utiliser pour faire des déclarations sur les objets qui vous intéressent – déclarations que vous pouvez ensuite traiter sur de grands jeux de données dans pratiquement toutes les langues.

La plate-forme est gérée par le Centre de recherche de Gotha et hébergée par l’ThULB Iéna. Elle s’adresse à des projets ayant un intérêt spécifique pour les données historiques.

En collaboration avec Wikimedia Allemagne et le GND de la Bibliothèque nationale allemande, nous essayons d’intégrer cette plateforme dans le prochain consortium d’instances fédérées de Wikibase comme ressource pour les « données de recherche ».

Pourquoi devrais-je utiliser FactGrid pour mes propres recherches ?

Le principal argument en faveur d’un compte FactGrid est la flexibilité imbattable du logiciel Wikibase, que nous avons réussi à installer, dans le cadre d’un projet pilote, en dehors de son site principal Wikidata et avec l’aide de Wikimedia Allemagne :

  • Vous recherchez un logiciel qui parle pratiquement toutes les langues – une plate-forme où vous pouvez entrer des données dans votre propre langue tout en permettant à d’autres de les lire dans leur propre langue ? Wikibase est ce logiciel.
  • Vous recherchez un logiciel qui vous permet de coordonner toute une équipe de manière transparente ? Dans Wikibase, c’est aussi simple que dans le logiciel MediaWiki de Wikipedia.
  • Vous recherchez un logiciel de base de données qui peut faire tout ce que les bases de données d’humanités numériques veulent normalement faire : analyses de réseau, représentations cartographiques, recherches croisées complexes, frises chronologiques (dans différents formats) – un logiciel qui se comporte presque comme un langage humain, tout en fournissant des services complets de base de données ? Wikibase est ce logiciel.
  • Vous avez des données provenant de projets antérieurs que vous voulez développer ? Wikibase dispose d’options de saisie automatique à grande échelle.
  • Vous voulez que vos données puissent être réutilisées ? Wikibase permet le téléchargement et le travail avec vos données, aussi bien hors ligne dans Excel qu’en dligne dans le cadre de nouveaux projets.
  • Vous voulez poser des questions entièrement nouvelles pour votre recherche ? Dans Wikibase, vous pouvez lier n’importe quel objet à n’importe quelle déclaration en fonction de vos intérêts.
  • Vous vous demandez ce qu’il adviendra de vos données et de vos outils de présentation une fois votre projet terminé ? Comptez sur une plate-forme sur laquelle vous ne travaillez pas seul et utilisez une licence de données qui permettra à d’autres personnes de continuer à travailler avec votre travail sans aucun risque !
  • Vous voulez poser des questions entièrement nouvelles dans votre recherche ? Dans Wikibase, vous pouvez relier n’importe quel type d’objet à n’importe quelle déclaration d’intérêt.
  • Vous vous inquiétez de ce qui arrivera à vos données et à vos présentations une fois votre financement terminé ? Comptez sur une plateforme où vous ne travaillez pas seul et utilisez une licence de données qui permet à d’autres de continuer à travailler à la fois avec vos données et vos outils !

Si vous recherchez une perspective à plus long terme, c’est ce que nous essayons d’offrir grâce à notre accord de collaboration en cours avec la Bibliothèque nationale allemande. Nous baserons notre plate-forme sur les données GND afin d’en faire aussi un outil grand public, avec pour objectif de devenir un acteur dans le paysage émergent des “installations fédérées Wikibase”.

Pourquoi ne pas utiliser Wikidata dès maintenant ?

C’est une question légitime à poser. Il y aura des projets (qui utilisent principalement des données) pour lesquels Wikidata sera la meilleure plateforme, comme l’Archivführer zur deutschen Kolonialzeit de la FH Potsdam. Reste que les projets Wikimedia (de même que GND) ne laissent pas de place à la recherche originale. Ils fonctionnent sur des “critères de notoriété” qui ne permettent pas la création à volonté d’objets et de relations entre objets innovants que les chercheurs voudraient tester.

Wikidata et le GND se concentrent sur les informations qui ont déjà été publiées ; ils mobilisent des travailleurs non chercheurs qui alimentent leurs bases de données à partir de recherches déjà publiées. Vous ne serez pas autorisé sur ces plateformes à énoncer des “hypothèses de travail” relevant de “votre recherche”. Vous ne pourrez pas créer des objets de base de données dans le seul but de mener sur eux à un stade ultérieur de votre recherche “rien de plus qu’une analyse statistique”.

Dans FactGrid, nous encourageons au contraire l’utilisation de la plate-forme comme un outil de recherche heuristique.

  • Créez des objets de base de données sur la plate-forme, quelle que soit par ailleurs leur pertinence pour une encyclopédie ou un catalogue de bibliothèque.
  • Risquez comme hypothèses de travail des chronologies provisoires en fonction de vos intérêts personnels.
  • Utilisez FactGrid afin de faire des déclarations non conventionnelles et qui n’ont d’intérêt que pour votre projet de recherche – le logiciel vous donne cette liberté.
  • Créez des objets de base de données spécifiques mentionnant votre projet de recherche dans les jeux de données que vous aurez substantiellement modifiés, ce qui vous permettra d’identifier votre contribution lorsque vous soumettrez votre recherche à votre institution de financement.
  • Risquez de nouvelles hypothèses sur la plateforme et indiquez votre point de vue par un numéro d’objet de base de données (servant en quelque sorte de “micro-publication”), afin d’attester de votre inventivité sur la base de données.

FactGrid est gratuit – comment est-ce possible ?

Le logiciel est disponible gratuitement et en cours de développement dans la communauté plus large des projets Wikimedia et au sein des institutions qui entendent utiliser Wikibase dans les prochaines années.

La plate-forme FactGrid est gérée par le Centre de recherche de Gotha sur un serveur virtuel de l’université d’Erfurt. L’URL allemande coûte 36 euros par an en frais de domaine, financés par le Centre de recherche de Gotha.

Tous les outils Wikidata sont à la disposition de nos utilisateurs. Ils comprennent toutes les applications standard dans les projets d’humanités numériques.

En outre, les logiciels et les outils étant open source, vous pouvez faire appel à votre prestataire de service informatique extérieur pour développer l’application spécifique dont vous auriez besoin.

Proposez à la communauté FactGrid vos propres développements d’outils et de présentation, ce sera le meilleur moyen pour que vos propres visualisations continuent à être développées après la fin du financement de votre projet. Si vous visez plutôt des solutions que vous voulez vendre, vous ne serez pas limité par la licence du logiciel. Vous pourrez commercialiser librement tout ce que vous aurez créé sur la base du logiciel ouvert.

Que dois-je faire des demandes de recherche non orthodoxes ?

Wikibase fait œuvre de pionnier dans la modélisation des données. Pour l’essentiel, vous ne créez que des relations entre des numéros Q (ou entre des numéros Q et des dates, des numéros Q et des coordonnées spatiales, des numéros Q et des fichiers média, des numéros Q et des URL).

Le logiciel ignore le type sémantique des relations que vous avez créées- il s’agit là encore de simples numéros P : Q1 – P1 – Q2 est un “triplet”, qui peut tout aussi bien signifier “Jean-Sébastien Bach (Q1) est le père de (P1) Carl Philipp Emanuel Bach (Q2) ” que “Cette lettre que j’ai trouvée dans les archives avec la cote XYZ (Q1) aurait été envoyée de (P1) Munich (Q2)”

Les numéros Q peuvent être attribués à toute espèce d’identité : personnes, documents, événements, idées… C’est vous qui décidez des types des numéros P dont vous avez besoin pour faire les déclarations qui vous intéressent. Vous n’avez pas à définir les objets dans un système de catégories a priori ; ce sont vos déclarations qui ajoutent de la chair aux objets que vous créez au fur et à mesure. Ne vous inquiétez pas si vous n’avez pas de modèle de données dès le premier jour. Faites des déclarations dès que vous en avez envie et voyez comment elles acquièrent la masse critique. C’est alors seulement que vous pourrez juger de la valeur de l’ensemble.

Toutes les déclarations peuvent elles-mêmes être « qualifiées » – “Jean-Sébastien Bach (Q1) était marié avec (P2) Maria Barbara Bach (Q2) à partir du (P2) 7 october 1707 (date) jusqu’à (P3) environ 5 juillet 1720 (date).” Toutes ces déclarations peuvent à leur tour être dotées de références : “cela ressort du (P4) registre paroissial de… (Q3)”, ”cela est indiqué dans (P5) la biographie bien connue de Bach XYZ (Q4)”.

Le système permet à tout moment de proposer des affirmations concurrentes. Elles sont simplement introduites avec leurs différentes sources et peuvent être comparées les unes avec les autres.

En fin de compte, n’importe quelle déclaration en langue naturelle peut être générée avec des triplets de ce genre, mais, surtout, cela permet d’exprimer cette déclaration dans n’importe quelle langue du monde : pour le système, toutes les déclarations ne sont que des liens entre des numéros Q et des numéros P. C’est vous seul qui attribuez aux numéros Q et P des “libellés” et des “définitions” qui leur donnent sens, et cela dans les langues avec lesquelles vous communiquez (le système gère par ailleurs les informations de date et de quantité dans toutes les normes mondiales avec une conversion automatique dans n’importe quelle direction) ; c’est là le secret des plateformes Wikibase, qui permet aux auteurs de saisir les informations dans leurs langues respectives et aux utilisateurs de lire ces informations dans n’importe quelle langue.

Quels sont les outils fournis par le système ?

Les entrées dans la base de données peuvent être effectuées une à une : ouvrez pour cela l’objet-id en question, allez au bas de la page de saisie et cliquez sur le lien “ajouter une déclaration”. Il vous sera alors demandé de saisir la déclaration que vous souhaitez faire. Vous n’avez pas besoin de connaître le numéro P. Il suffit d’indiquer l’objet dans la langue que vous utilisez et de cliquer sur l’auto-complétion qui vous est proposée. La plate-forme utilisera pour vous le numéro P de cette déclaration. Indiquez alors dans le champ qui s’ouvre l’objet de votre déclaration. Le système, là encore, vous proposera, à mesure que vous tapez le texte de votre déclaration, des suggestions de plus en plus précises.

Les entrées dans la base de données peuvent également être créées et enregistrées automatiquement à partir de tableaux Excel ou CSV. (Ceci est le masque de saisie et voici le guide succinct pour le faire).

Les requêtes dans la base de données doivent être formulées sous forme d’interrogations “SPARQL”, un langage de requêtes qui n’est (malheureusement) pas si facile à utiliser, mais qui, au fond, n’est pas plus complexe que les recherches que vous pourriez vouloir effectuer.

Le plus souvent les utilisateurs de SPARQL ne savent pas écrire leurs requêtes dans le code source. Vous pouvez cependant utiliser des modèles de requêtes où sont indiquées les entrées qu’il faut modifier afin d’exécuter votre recherche spécifique.

En outre, si vous savez exactement le type de requêtes que vos utilisateurs doivent exécuter, vous pouvez créer vos propres masques de saisie, comme ceux que vous utilisez dans les interfaces habituelles des bibliothèques en ligne, qui parleront alors SPARQL avec la base de données.

Le système comprend aussi des outils cartographiques, des frises chronologiques, des réseaux, des arbres généalogiques, des graphiques, etc. Vous n’avez pas besoin de télécharger des applications particulières. Vous demanderez à SPARQL de produire la représentation que vous essayez d’obtenir. Le projet Scholia sur Wikidata présente certaines de ces visualisations.

Que dois-je faire si je veux donner mes représentations de données sur ma propre plate-forme ?

Cela ne devrait pas poser de problème technique. Uwe Jung a montré comment l’interface de la FH Potsdam utilise Wikidata comme dépôt de données sans laisser les utilisateurs voir la base de données à laquelle ils accèdent.

Il n’y a rien de mal à utiliser FactGrid comme dépôt externe et à monter son propre projet de recherche sur le serveur de son université d’origine, en y proposant des accès ciblés à la base de données selon un modèle de recherche de son choix.

FactGrid octroie essentiellement des licences d’utilisation des données à CC0 – cela veut-il dire que je renonce à tous les droits sur mes recherches ?

Opter pour la licence Creative Commons signifie essentiellement que vous conservez tous les droits sur l’ensemble de vos données. Mais surtout, la licence CC0 signifie que vos données deviennent librement utilisables et que vous pouvez ainsi réduire le danger de recherches obsolètes à long terme.

Quelques considérations de base : CC BY 4.0 est à première vue la licence que les scientifiques préféreront. Elle permet l’utilisation gratuite des données à la condition que celles-ci soient correctement citées. En pratique, cela fonctionne pour les textes (comme ce billet de blog) ; dans ce cas, on voit clairement comment on aimerait que le texte soit cité : avec une référence à l’auteur, le titre de la publication, le lieu de publication et la date. Mais supposons que vous souhaitiez que vos données soient citées, disons dans une visualisation ? Une lettre envoyée de Paris à Berlin en juin 1753 se réduisant à une ligne sur une carte, comment cette ligne doit-elle être correctement annotée ? Comment voulez-vous être cité si vous n’avez fait qu’améliorer un ensemble de données existantes ? Les licences “share alike” sont encore plus problématiques : “Ces données sont disponibles gratuitement si les utilisateurs ultérieurs les gardent tout aussi libres”. Cela semble être le plaidoyer ultime pour la gratuité. Mais comment un sous-utilisateur peut-il s’assurer que ses sous-utilisateurs respecteront à leur tour votre contrat de licence (surtout si ce sous-utilisateur offre ses données sous CC0) ? Les sous-utilisateurs seront bien avisés de ne pas utiliser de données provenant de plateformes CC-BY ou CC Share-Alike.

Nos entreprises communes avec Wikidata et la Bibliothèque nationale allemande ne nous ont finalement laissé qu’une seule option : rendre nos données aussi librement disponibles que nos partenaires, autrement dit sous CC0, c’est-à-dire sans garantie que les utilisateurs ultérieurs préciseront toujours exactement qui a collecté les données, ni sur ce que les utilisateurs tiers seront autorisés à faire avec ces données.

 
En pratique, la licence ouverte maximale ne signifie pas que les données de FactGrid sont des données sans auteur, bien au contraire. Outre que nous suggérons aux utilisateurs de toujours citer la recherche qu’ils utilisent, nous faisons l’hypothèse que Wikidata et le GND souhaiteront renvoyer à la recherche sur notre plate-forme. Toutes les modifications apportées aux jeux de données sont liées par le système à nos vrais noms, visibles par tous les utilisateurs. Si un jeu de données a été tout particulièrement travaillé par un projet de recherche, vous pouvez l’indiquer dans une note à part qui sera transférée avec ce jeu de données. Vous pouvez également noter votre travail dans le jeu de données lui-même. Enfin, tout le monde peut interroger la base de données pour savoir quels jeux de données ont été travaillés dans le cadre d’un projet particulier.

En fait, les bases de données comme Wikidata ou le GND de DNB sont intéressées à citer la recherche – cela renforce la solidité de leurs données, et FactGrid est dans la position unique de fournir aux deux institutions une plate-forme sur laquelle les gens peuvent faire ce qu’ils ne pourraient pas faire sur leurs grandes plates-formes.

Que se passe-t-il si je veux continuer à travailler avec mes données sur une autre plateforme ?

Puisque vous avez saisi vos données sans restriction de droits d’auteur, vous pouvez travailler librement avec elles sur tout autre projet qui vous intéresse. En fait, nous aimons être “juste un incubateur” pour les données de recherche.

Que se passe-t-il si les utilisateurs de FactGrid se disputent sur une date “correcte” ?

Le logiciel permet de traiter des données contradictoires – ce qui est particulièrement intéressant dans le domaine de la recherche historique où nous disposons souvent de preuves documentaires contradictoires, sans pouvoir être certain de l’information correcte. Les noms sont traités avec des orthographes différentes ; il arrive que les historiens se contredisent.

Le système permet de reproduire la situation contradictoire ; il permet d’étayer les déclarations avec des dizaines de références et dans différentes orthographes.

Des déclarations divergentes peuvent être comparées les unes avec les autres – par exemple, la déclaration qui fait actuellement autorité et les variantes qui ne circulent qu’en raison des diverses sources contradictoires.

Que deux chercheurs arrivent à des résultats différents, voilà qui devrait en général vous intéresser. Le danger majeur est d’avoir fait une hypothèse erronée et qu’un autre projet sur une autre plateforme donne la solution de l’énigme et travaille sur la bonne date sans que vous le sachiez, ou pire, sans que vous ayez même la possibilité de corriger votre erreur des années après la fin de votre projet.

Pourquoi devrais-je risquer la transparence de mon projet dès le début ?

C’est probablement le problème le plus difficile, celui qui empêche actuellement certains projets d’utiliser la ressource que nous avons ouverte. L’alternative est une ressource accessible seulement avec un mot de passe aux membres de l’équipe jusqu’à la date de publication, c’est-à-dire quasiment jusqu’à la fin du projet. Aucun projet concurrent ne peut alors s’emparer des résultats, du moins en théorie. Personne ne peut voir l’erreur par où vous avez commencé et que vous avez ultérieurement corrigée. Personne ne peut voir non plus le travail fourni par les assistants qui entrent les données dans la base, ni l’implication réelle du chef de projet – tels sont les avantages supposés d’un travail non transparent sur une plateforme qui ne sera mise en ligne qu’à la fin de votre financement.

La recherche transparente offre ses propres garanties : si vous trouvez un document révolutionnaire et établissez une connexion décisive, c’est l’occasion d’attacher la découverte à votre nom et à votre projet. Si quelqu’un, demain, fait la même découverte dans les archives que vous venez de visiter, pas de chance pour lui : vous aurez enregistré votre observation avec un lien dans l’historique des versions que vos rivaux ne pourront pas nier.

En même temps, la plate-forme collective invite à coopérer. Expliquez clairement aux autres équipes sur quoi vous travaillez et permettez-leur de vous contacter sur la plate-forme !

Les inconvénients d’un site web prétendument sécurisé et qui ne sera mis en ligne qu’à la fin du financement du projet sont sérieux : Lorsque le projet est publié, le temps des échanges avec les utilisateurs est déjà révolu. Si la mise sur Internet se fait dans les dernières semaines, celles où le projet est sous pression, vous vous trouverez totalement incapable de réagir par des changements plus conceptuels. Et si vous avez mené des recherches uniquement pour la publication d’un livre, que ferez-vous, vous et votre équipe, des données que vous avez rassemblées dans des fichiers Word et des feuilles de calcul Excel ? Personne ne pourra les verser dans des bases de données, car l’harmonisation à ce stade tardif sera un obstacle insurmontable. Votre seul espoir est que des lecteurs de votre livre parcourront toutes vos notes de bas de page pour en tirer des corrections pour nos catalogues de bibliothèque et pour différents projets Wikipédia. Le risque, finalement, est d’avoir un livre sans impact sur la base de données collective et sur les projets d’humanités numériques et qui sera, à cet égard au moins, obsolète dès sa publication.

L’avenir devrait résider dans une nouvelle attitude à l’égard de la base de données publique. Les chercheurs devraient pouvoir corriger et élargir encore cette base chaque fois qu’ils y accèdent. Pour cela, il leur faut une incitation et une sécurité que seul peut leur donner un environnement de recherche dans lequel le travail soit référençable et citable. Pour cela, Wikibase est mieux équipée que tout autre système.

Comment puis-je faire accepter mon projet sur FactGrid ?

La plate-forme FactGrid ne comporte pas de couche profonde invisible. Tout le monde peut interroger la base de données et les requêtes donneront les mêmes informations, que vous soyez connecté ou non. Votre compte d’utilisateur personnel présente simplement l’avantage de vous permettre de passer à votre langue préférée lorsque vous consultez les données et de voir le lien d’édition sur chaque déclaration.

Si vous souhaitez alimenter la plate-forme avec vos propres données et si vous souhaitez y mener un projet, vous devez disposer d’un compte. Les comptes sont donnés sous des noms réels par les administrateurs. Le logiciel fournit un lien “demande de compte”. Vous pouvez également nous contacter par courrier électronique. Les chefs de projet peuvent recevoir des comptes administratifs leur permettant de donner accès aux membres de leur équipe et aux utilisateurs qui les intéressent.

Une fois connecté, vous pouvez saisir des données en masse ou apporter des corrections spécifiques où bon vous semble. Toute entrée sera connectée à votre compte d’utilisateur. D’autres utilisateurs peuvent annuler vos modifications, mais non sans laisser une trace documentée de cette intrusion dans l’historique des versions – visible par le monde entier.

Si vous souhaitez travailler sur un projet plus complexe, qu’il s’agisse d’une recherche familiale personnelle, d’une visualisation unique dont vous auriez besoin pour une communication, ou encore de l’intégration dans la base de milliers de documents que vous auriez rassemblés dans le cadre d’un projet de recherche de 5 ans, parlez-en à ceux qui sont déjà sur FactGrid et à ceux qui organisent la plate-forme. Nous ne serons pas (nécessairement) désireux de signer un protocole d’entente avec vous, mais il pourrait être très intéressant de faire connaître votre projet sur le blog, ainsi que sur l’ensemble de la plateforme. Là où votre travail devient passionnant, c’est lorsque vous modifiez le travail que d’autres ont déjà fait et que vous encouragez les acteurs d’autres projets à adopter les bons modèles que vous introduisez. Il n’est pas indispensable de discuter des modèles de données avec tous les autres utilisateurs, mais cela peut aider, ne serait-ce que pour diffuser votre travail sur la plateforme. Adoptez des requêtes de recherche composées par d’autres, découvrez des visualisations auxquelles vous n’avez pas pensé, obtenez de l’aide sur la plateforme.

FactGridest conçu pour gérer un environnement de recherche excitant que vous ne trouverez pas ailleurs.

traduit par Bruno Belhoste

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).
Jack Kirby, “The Fourth Dimension is a many splattered thing!” from Alarming Tales, 1 (September 1957).

The Illuminati Correspondence Fast Forward

Paul-Olivier Dehaye scripted this visualisation for us (using Uber’s http://Kepler.gl). An html-file that captures all the Illuminati exchanges from the 1770s into the 1790s as far as we have spotted them (there are some misfits in this visualisation which we can now suddenly identify and which need to be eliminated on the database; the visualisation itself is basically a screenshot, it does not adapt to changes in the database).


<click to play>

The idea to represent letters in lines on a map together with a timeline on which the user can set a span that can then be shifted through the timeline – has become a classic in recent years: The Stanford Republic of Letters project seems to have been the first to come up with this visualisation ten years ago

Nodegoat is offering this visualisation as a standard aplication.

The visualisation is cool for correspondences since letters happen to travel on maps from senders to recipients (or to multiple recipients as soon as letters are forwarded – a standard procedure in all hierarchical Illuminati exchanges).

The simple visualisation which the SPARQL query service had yielded was already interesting to look at:

Illuminati Letters – sent from where?

It showed the sender’s places of all known Illuminati letters and vaguely hinted at the Illuminati centres in Germany. But the picture remained static. It lacked the directions and the historical drama. You got more information if you clicked at an individual dot – usually this would open just a speech bubble with information about the particular letter that created this dot. But here and there one would see far more: a dot sparkling a fireworks of dots as in the case of Weimar (from where Christoph Bode organised the Order as the de facto leader after 1785/86). The beautiful bouquet is otherwise misleading – the dots do reach out to the various destinations; they stand for quantities which you only see if you hit the right dot (and which then obliterate much of the rest of the picture).

Bode’s Weimar correspondence – a nice representation of the number of documents but not much more.

Letters – that is the charm of the more refined temporospatial visualisation – tend to come in correspondences and these evolve, they stretch out, they blossom, and they die eventually.

The Illuminati are an almost ideal object for this particular visualisation as they present a full case to study. The “Republic of Letters” had remained out of reach for the Stanford project. The thing which we today prefer to call “academia” or “scientific community” was far bigger than the carefully selected and spectacular cases that created the first visualisations in 2009 — and only the whole picture would have revealed the evolution of our present academic debates between the 1480s and 1800. Regions and emerging nations brought forth increasingly scattered cultures of learning and they invented modern national topics such as our present debate of (usually national) literature. The spectacular exchanges could not possibly reveal these developments.

The Illuminati correspondence is, admittedly, no longer complete. We are looking here basically at the archives of Weishaupt and Bode and the Bavarian publications of 1787 – but it is even in this selection a clearly and well defined object. You know when you have an Illuminati letter in front of you: The author uses code names and the (messy) Illuminati-Persian calendar. The Organisation created complete genres of letters such as the monthly “Quibus Licet” which every member had to hand in and which would be answered with a “Reproche” signed by “Basilius” two months later. The genres are as remarkable as the internal affairs discussed in these letters.

The Order itself was at the same moment basically a complex correspondence: an organisational construct that generated a particular and unstable flow of information. Paul’s time lapse encapsulates the history and the drama of the Illuminati: For about two years – from 1776 to 1778 – we see very little: Ingolstadt is the place where Weishaupt’s “Perfectibilists” could organise most of their affairs in face to face meetings.

The situation changed dramatically in 1778: The Illuminati opened a lodge in Munich and started to infiltrate the masonic world.

Again two years later, in 1780, we see Adolph Freiherr von Knigge rising with exchanges he is now maintaining from Frankfurt and then from Heidelberg.

Knigge wins Bode for the Order in September 1782 and Bode in turn resolves the escalating conflict between Knigge and Weishaupt in 1784 and 1785: Knigge is forced to withdraw but Weishaupt does not regain his former position as the head of the organisation. Bavaria exposes the Order in 1786/87 with the first two editions of intercepted Illuminati documents. Weishaupt flees to Regensburg and then to Gotha where he ends in personal ignominy while it remains Bode’s part to come to the conclusion that he would not reform the secret organisation which could no longer claim to be a secret society.

You can pinpoint the individual letter to see who was writing here to whom with what letter exactly

Paul’s abstract movie wants to be explored. You can define the time frame and push it manually through the timeline, and you can click at each line to see whose letter is creating it. We can now see the overall quantities and the processes, the organisation’s actual growth and collapse on the map.

Far from perfect

The visualisation is state of the art and yet not much more than a show case at the moment. It is neither created in ever fresh queries nor can you use the html-page without some coding for your own questions.

The message, however is clear: The interface one would love to have with this visualisation could be far more simple than the present SPARQL query service. One would design it to always explore “correspondences” (P122Q11243) and one would ask the user for simple P—Q specifications of his or her desired visualistion. In the Illuminati case this would be:

  • Research Interest (P97) — Illuminati (Q10677)

One might just as well ask for “author” and name a couple of authors, or for recipients, but the Interface would do the rest and run the SPARQL query of correspondences to generate a list of the letters in question with senders’ and receivers’ places and dates.

Paul’s message is that this can be done: The centre of his html file is basically a SPARQL-search from our database. Maybe someone will script the interface during the Wikidata.con next month in Berlin.

And one would love to have more…

  • Think of a visualisation that tracked the movements of people on maps and that captured the moments when people (could have) met.
  • Think of a visualisation of the dissemination of objects – like copies of a specific edition.
  • Think of the spread of an organisation like Freemasonry with its mother lodges and filial branches.
  • Imagine a visualisation that traced your ancestors with a look at genealogical data.

Wikibase instances should invite the production of interfaces that can do certain jobs without bothering the user with SPARQL or Java proficiency. The cool thing about Wikidata or FactGrid is that these instances develop their (more or less) static ways to organise core data. We get more data but we continue to make the same useful statements especially if we have visualisations asking for these statements to be made with certain P- and Q-numbers.

Nodegoat would not lose any its charm if we began to offer similar visualisations; it will remain the software for the vast majority of projects that prefer the exclusive environment on which they can run exclusive and definitive presentations of their data. Wikibase instances are already very different beasts: They invite projects that want an open and growing landscape of data, and these projects will show the far bigger need of standard visualisations (i.e. of visualisations that use existing data and existing data structures). Projects on FactGrid or Wikdata need interfaces that already speak SPARQL. Time then to look at the more complex visualisations that are already running in environments such as the Stanford Republic of Letters Project or the Cultures of Knowledge Project in Oxford. We should act as the coolest competitors in this field, gather inspiration with the aim to inspire in turn.

FactGrid FAQ – Why should I use FactGrid for my research project?

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).

auf Deutsch
en français
magyarul

What is FactGrid?

FactGrid is a Wikibase installation — that is is both, a regular wiki and a database which you can use to make statements about objects of your interest — statements which you can then handle in practically any language in big data sets.

The platform is run by the Gotha Research Centre and hosted by the ThULB Jena. It addresses projects with a specific interest in historical data and is part of the German National Research Data Infrastructure NFDI4Memory.

In joint ventures with Wikimedia Germany and the German National Library’s GND we are trying to bring this platform into the upcoming consort of federated Wikibase instances as a resource for research data.

Why should I use FactGrid for my own research?

The biggest argument for a FactGrid account is the unbeatably flexible software, Wikibase, which we have managed to implement in a pilot project with the help of Wikimedia Germany – outside its primary location, Wikidata:

  • You are looking for a software that speaks practically any language — a platform on which you can enter data in your language and allow others to read your data in their languages? Wikibase is this software.
  • You are looking for software in which you can transparently coordinate a whole team? In Wikibase this is as easy as in the Wikipedia software MediaWiki .
  • You are looking for a database software that can do everything Digital Humanities databases normally want to do: network analysis, map representations, complex linked searches, timeline representations (in various formats) – a software that almost acts like human language yet provides full database services? Wikibase is this software.
  • You have data from previous projects which you want to build on? Wikibase has large-scale automatic input options.
  • You want to make sure that that other projects will actually use your data? Use a platform that allows the download and further work with your data offline in Excel or online in ever new projects.
  • You want to ask entirely new questions in your research? In Wikibase you can link any sort of objects with any kind of statements of your interest.
  • You are worried what will happen to your data and presentations once your funding is over? Work on a platform where you do not stay alone and where, thanks to the CC0 license you are using, you encourage colleagues to continue right were you stopped!

If you are looking for a long term perspective this is what we are trying to offer in our present joint venture with the German National Library. We will base our platform on GND data in order to serve as a broad public tool and with the aim to become a player in the emerging landscape of “Federated Wikibase installations”.

Why not use Wikidata right away?

That is a legitimate question to ask. There will be projects (projects that are mainly using data) for which Wikidata will be the better platform. The FH Potsdam’s “Archivführer zur deutschen Kolonialzeit” has demonstrated the beauty of working directly on Wikidata; we talked about this with Uwe Jung, who demonstrated the technical solutions they have chosen in Potsdam.

On the other hand, there remain basically two things which you will not be able to do on Wikidata or on a platform like the GND: Wikimedia projects (and the GND) have strict “No original Research” policies and observe fundamental decisions to operate on “criteria of notability“, which will not allow the arbitrary opening of database objects and the innovative object relationships researchers would like to test.

Wikidata and the GND focus on information that has already been published and on non-research workers who feed the respective databases from published research. You will not be allowed to state a “working hypotheses” of “your research” on these platforms. You will not be able to create entities with the aim to run “nothing but a statistical analysis” on them at a far later stage of your work.

In FactGrid we encourage the use of the platform as a heuristic research tool.

  • Create database objects on the platform, no matter what their relevance in an encyclopedia or in library catalog could be.
  • Risk provisional chronologies as working hypotheses along with your personal assumptions.
  • Use FactGrid in order to make unconventional statements that are presently interesting only in your research project – the software gives you this freedom.
  • Create specific database objects that state your research in all the data sets which you have substantially modified and become able to submit your research with the particular item as the envelope to your funding institution.
  • Risk any new thesis on the platform and state your respective view with a database object number as a “micro-publication” in order to claim your ingenuity on the data base.

FactGrid is free – how does it work?

The software is freely available and under development in the larger community of Wikimedia projects and among the institutions which are going to use Wikibase over the next years.

The FactGrid platform is presented by the Gotha Research Centre on a virtual server of Jena’s University Library free of charge.

All the Wikidata-tools are available to our users. These include all the standard applications of regular Digital Humanities projects.

With both software and tools being open source you can use any favorite software company you are working with, to generate the specific application which you feel you need.

If you feed your tools into the open compound this will be your best way to make sure that future projects will continue their development.

If you aim at technical solutions which you want to sell with financial profit that again will not be restricted by the software license. You can freely commercialize whatever you create on the basis of the open software.

What do I do with unorthodox research interests?

Wikibase is groundbreaking in its data modelling. Essentially you are only creating relations between Q-numbers (or relations between Q-numbers and dates, Q-numbers and space coordinates, Q-numbers and media files, Q-numbers and URLs).

The software does not know what kind of relationships you are stating – these again are just P-numbers: Q1 – P1 – Q2 is a “triple” and can mean “Johann Sebastian Bach (Q1) is the father of (P1) Carl Philipp Emanuel Bach (Q2)”; it can mean just as well “This letter which I found in the archive with the shelf mark XYZ (Q1) has allegedly been sent from (P1) Munich (Q2)”

Q-numbers can be assigned to anything imaginable – people, documents, events, ideas… You decide what kinds of P-numbers you need in order to make statements of your interest. You do not define objects in a system of fixed categories; your statements are adding colour and solidity to whatever object you create as you go along. Do not worry if you do not have the data model on day one. Make statements when you suddenly want to make them and see how they gain the critical mass that can eventually be evaluated.

All statements can be “qualified” – “Johann Sebastian Bach (Q1) was married to (P2) Maria Barbara Bach (Q2) beginning on (P2) October 7, 1707 (date) ending (P3) about July 5 1720 (date).” All of these statements can in turn be equipped with references : “this is clear from (P4) the church book of… (Q3)”,”this is stated in (P5) the well known Bach biography XYZ (Q4)”.

The system allows competing claims at any time. They are simply introduced with their different sources and can be balanced against each other.

You can create practically any normal language statement with triples of this depth of specificity; but above all, this opens the door to the world of statements in all the various languages you might want to speak: The system operates with Q- and P-numbers; the rest is labels in languages which you want to offer to your users. The software will in addition translate dates and quantities into other formats; this is basically the secret that makes it possible for Wikibase platforms to be edited by people in their own languages and to be read by the world in practically any other language.

Which tools does the software provide?

Database entries can be made one by one: Open the object-id in question, go to the bottom of the input page and click the “add statement” link. You will now be asked for the statement you want to make. You do not need to know the P-number. State the property in the language you are using and click at the auto complete you are eventually being offered. The platform will now use the P-number of that statement for you. Complete in the next box that opens your statement. You will again get suggestions to use as you are typing.

Database entries can also be created and substantiated in automated inputs from Excel or CSV lists. (This is the input mask and this is the short guide to it.)

Database queries have to be formulated as “SPARQL” queries, a search language that is (unfortunately) not that easy to use, but that is eventually as complex as the searches you might want to perform.

SPARQL users do not necessarily know how to write SPARQL source code. You usually use sample queries that tell you where you have to change the input in order to run your particular search.

If you know exactly which type of search queries your users should run you can create your own input masks just as you know them from conventional online library interfaces, which will then speak SPARQL with the database.

The software package includes illustrations on maps, timelines, networks, genealogical relationships, graphs, and so on. You do not need to download particular applications. You will ask SPARQL to produce the representation you are trying to get. The Scholia project on Wikidata has a beautiful first presentation of some of the visualisations.

What do I do if I want to give my very own data representations on my own platform?

That should not be a technical problem. Uwe Jung demonstrated how the FH Potsdam interface uses Wikidata as its data repository, without letting users see the database they are accessing.

There is nothing wrong with using FactGrid as an external repository and building up your own research project on the server of your home university, where you can offer targeted database accesses under a typical search template of your choice.

FactGrid basically licenses data to CC0 – does that not mean that I give up all rights to my research?

Opting for the Creative Commons 0 license means essentially that you continue to be free to do whatever you want with your data – you, and not your publisher or the platform that received your data under a scheme, continue to control. But above all, the CC0 license means that your data becomes freely usable and that you can thus reduce the danger of obsolete research in the longer run – others will continue to root out the ugly mistakes you could not hope to correct.

Some basic considerations: CC BY 4.0 is at first glance the license that scientists will prefer: It allows the free further use as long as it receives the accurate citation. In practice, this will work for texts (such as this blog post); here it is clear how one would like to see the text cited: with a reference to one’s own name, with the title of the publication, the place of publication and the date. But do you want your data quoted let us say in a visualisation? A letter sent from Paris to Berlin in June 1753 will be a line on a map and how should this line be properly annotated? How do you want to be quoted if you only improved a data set? “Share alike” licenses are even more problematic: “These data are freely available if the subsequent users keep it just as freely available.” That sounds like the ultimate plea for free use. But how can a sub-user ensure that his sub-users, in turn, will respect your license agreement (especially if this sub-user is offering his data on CC0)? Sub-users are well advised not to use data from CC-BY or CC Share-Alike platforms.

Our joint ventures with Wikidata and the German National Library left us only only one option: to make our data as freely available as our partners: CC0, that is without ensuring that subsequent users will still specify exactly who collected the data, and what third-party users are allowed to do with that data.


In practice, the maximum open license does not mean that FactGrid data is data without authorship, quite the contrary. We suggest that research is cited and assume that Wikidata and the GND are only too happy to state research from our platform. All changes are linked to respective the author names. If research projects have worked more substantially on a data set, they will have stated this in a separate note on the data set that can now be adopted with the data transfer.

Databases like Wikidata or DNB’s GND are in fact interested to quote research – it boosts their data solidity, and FactGrid is here in the unique position to give both institutions a platform on which people can do what they cannot do on the respective larger platforms.

What happens if I want to continue working with my data on another platform?

Since you entered your data without a copyright restriction, you are free work with them on any other project of your interest. We actually like to be “just an incubator” for research data.

What happens when FactGrid users argue about a “correct” date?

The software makes it possible to handle contradictory data. This is particularly interesting in the field of historical research, where we often have conflicting documentary evidence without being able to determine the correct statement after so much time. The software makes it possible to reproduce such a contradictory situation. Numerous statements can be given side by side with their various respective sources. You can then still balance the statements against each other – either by turning one of them into the statement to privilege in future searches and/or by adding qualifying statements with your personal evaluations.

If two researchers come to different results, think of it as the situation you should actually be interested in. It is far worse that you have made mistakes on a platform where they will never be corrected and where they eventually discredit all your work as obsolete beyond repair.

Why should I risk transparency in my project right from the start?

This is likely to be the toughest issue that currently prevents projects from using the resource which we have opened. The alternative is the resource, which is accessible to the team only under passwords until the publication deadline is reached almost at the end of the project. No competing project can snatch away findings, so the theory. No one sees where you have initially made a mistake. Nobody records what assistants are typing in and where the project leader is involved – such are the presumed advantages of non-transparent working on a platform that will only go online at the end of your funding.

Transparent research offers its own securities: If you find a groundbreaking document and establish a decisive connection, then this is your chance to fix the statement to your name and project. If someone makes the same discovery tomorrow in the archive you have just visited, tough luck. You will have recorded your observation with a link in the version history which your rivals will not be able to deny.

At the same time, the collective platform expresses the invitation to cooperate. Make it clear to other teams what you are working on and allow them to contact you on the platform!

The risks of the allegedly secure website, which is only going online at the end of the project’s funding, are serious: The time for an exchange with users is over. The internet presence goes online in the heated final weeks while the project is totally unable to react with more conceptual changes. If you have done research solely for a book publication, it will remain unclear what you and your team should do with the data, you have still collected in Word files, and Excel spreadsheets. Nobody will be able to feed all these data into any resource – the harmonisation at this late stage will be an insurmountable obstacle. You can only hope that readers of your book will scan all your footnotes for corrections that should reach our library catalogs and the various Wikipedia projects. The risk is here the book that has no influence on the collective data base and of DH projects that become obsolete right after their publication.

The future should lie in a new attitude toward the public data base. Researchers should be able to correct and to further widen this base wherever they access it. The incentive and the security they will need here is the research environment in which they can mark their work and make it citable. For this Wikibase is better equipped than any other software.

How do I get my project accepted on FactGrid?

The FactGrid platform has no invisible deeper layer. Anyone can query the database and the queries will give the same information whether you are logged in or not. Your personal user account just has the advantage that you can now switch to your favorite language when looking at data, and that you see the edit link on each statement.

If you want to feed your own data into the platform and if you want to run a project on the the platform, you will need an account. These are given under real names by the administrators. The software provides an “account request” link. You can also contact us via email. Project leaders can receive administrative accounts which they use to assign to team members and users of their interest.

Once logged in, you can enter data in bulk or make specific corrections wherever you feel like. Any input will be connected to your user account. Others can undo your edits but not without leaving a documented mark of that intrusion in the version history – visible to all the world.

If you want to work on a more complex project —

  • that can be personal family research,
  • it can be a single visualization you need in a seminar paper,
  • it may just as well be the input of thousands of records in a 5-year research project

— speak with those on board and those organising the platform. We will not (necessarily) be interested to sign a public memorandum of understanding with you but it might be cool to advertise your project on the blog and to make it known on the entire platform. Your work becomes exciting, where you modify the work others have already done and where you encourages players from other projects to adopt good models which you are introducing. You do not need to discuss data models with all the others but using models with all the others is also a way to spread your work and to make it appear in queries composed by others, and visualizations you did not think of.

The software is designed to manage both: unique statements which only you are interested in and statements that will spread far beyond your own initial research interest as you are now feeding the unforeseen queries which others will run on their and your data.

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).
Jack Kirby, “The Fourth Dimension is a many splattered thing!” from Alarming Tales, 1 (September 1957).

FactGrid FAQ – Warum sollte ich das FactGrid für die eigene Forschung nutzen?

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).

in English
en français
magyar nyelven

  1. Was ist das FactGrid?
  2. Warum sollte ich das FactGrid für die eigene Forschung benutzen?
  3. Warum nicht gleich Wikidata benutzen?
  4. Das FactGrid kommt kostenfrei – wie geht das an?
  5. Was mache ich mit unorthodoxen Forschungsanliegen?
  6. Welche Tools stellt die Software mir zur Verfügung
  7. Was mache ich, wenn ich ganz eigene Datendarstellungen auf meiner eigenen Plattform realisieren will?
  8. Das FactGrid lizenziert Daten grundsätzlich CC0 – heißt das nicht, dass ich alle Rechte an meiner Forschung aufgebe?
  9. Was geschieht, wenn ich mit meinen Daten auf einer anderen Plattform weiterarbeiten will?
  10. Was passiert, wenn es zwischen FactGrid-Benutzern zum Streit über ein “korrektes” Datum kommt?
  11. Warum sollte ich in meinem Projekt Transparenz schon von Anfang an riskieren?
  12. Wie nutze ich das FactGrid konkret?

Was ist das FactGrid?

Das FactGrid ist eine Wikibase-Instanz, die, vom Forschungszentrum Gotha aus organisiert, an der ThULB Jena gehostet wird.

Wikibase ist die Software, die hinter Wikidata läuft. Die Instanz will historischer Forschung einen eigenen Freiraum zur Verfügung stellen und im Verlauf technisch in der Lage sein, mit der GND und mit Wikidata laufend Daten auszutauschen — Daten, die aus dem FactGrid heraus international als Forschungsdaten zitierbar werden. Seit dem 5. November 2022 ist das FactGrid — als global agierende Plattform — offizielles Repositorium im Spektrum der deutschen Nationalen Forschungsinfrastruktur, NFDI4Memory.

Sehen Sie hierzu auch den Beitrag zu den Zielerwägungen, den Barbara Fischer für die GND und Jens Ohlig für Wikimedia am 9. Mai 2019 im Wikimedia Blog veröffentlichten.

Warum sollte ich das FactGrid für die eigene Forschung benutzen?

Dafür spricht erstens die unschlagbar flexible Software, Wikibase, die wir im Pilotprojekt mit Wikimedia Deutschland bahnbrechend außerhalb ihres eigentlichen Orts, Wikidata, zum Laufen brachten:

  • Sie suchen eine Software, die praktisch jede gängige Sprache spricht und in der sich Daten in jeder Sprache eingeben und in danach in beliebigen anderen Sprachen ausgeben lassen? Wikibase ist diese Software.
  • Sie suchen eine Software, in der Sie ein ganzes Team transparent koordinieren können? In Wikibase ist das so einfach wie in der Wikipedia Software MediaWiki.
  • Sie suchen eine Datenbank-Software, die alles kann, was Datenbanken der Digital Humanities normalerweise können: Netzwerkanalysen, Repräsentationen auf Landkarten, komplex verknüpfte Recherchen, Timeline-Darstellungen (in verschiedenen Datenformaten), und die sich dabei einem Umgang mit normalen Aussagen annähert? Wikibase ist diese Software.
  • Sie haben Datenmengen aus vorheriger Forschung, auf die Sie aufbauen wollen? Wikibase erlaubt den großflächigen automatischen Input.
  • Sie wollen sicherstellen, dass Ihre Daten nachnutzbar werden? Wikibase ist genau darauf eingestellt.
  • Sie wollen neuartige Fragen an Material stellen? In Wikibase können Sie beliebige Objekte mit beliebigen Aussagen Ihres Interesses verknüpfen.
  • Sie fragen sich, was nach Ihrer Projektlaufzeit mit Ihren Daten und Ihren Präsentationstools noch geschieht? Setzen Sie auf eine Plattform, auf der Sie nicht allein arbeiten und auf eine Datenlizenz, die es anderen ermöglicht, Ihre Arbeit wirklich risikolos zu nutzen und fortzuentwickeln!

Für das FactGrid spricht im selben Moment, dass wir im Interesse am breiten und offen bleibenden Datenaustausch in einer Kooperation mit der Deutschen Nationalbibliothek die gesamte Plattform soeben auf GND-Daten aufsetzen, um das Projekt in seiner größeren Breite in der entstehenden Landschaft von “Federated Wikibase Platforms” zu verankern.

Warum nicht gleich Wikidata benutzen?

Das ist eine gute Frage, die man sich stellen sollte. Es wird Projekte geben (die hauptsächlich Daten nutzen), für die Wikidata die bessere Plattform ist. Für die FH-Potsdam demonstrierte dies der „Archivführer zur deutschen Kolonialzeit“; wir sprachen darüber mit Uwe Jung, der die technischen Lösungen vorstellte, die man dort realisierte.

Zwei Dinge, gibt es, die es andererseits interessant machen, gezielt eine (Wikidata und die GND beliefernde) eigene Plattform aufzubauen: Die Wikimedia- (und GND–) spezifische “No Original Research” Grundregel und die fundamentale Entscheidung für Relevanzkriterien auf beiden Plattformen, die das beliebige Aufmachen von Datenbankobjekten und Objektbeziehungen ausschließt.

Wikidata und die GND richten sich auf bereits publizierte Information aus und auf Bearbeiter außerhalb der Forschung, die Publikationen auswerten und Daten eingeben. Unmöglich wird es bleiben, auf diesen Plattformen Arbeitshypothesen zu riskieren, Datierungen, die erst einmal auf Probe gestellt sind, Personeninformationen, die später in einer Publikation nur statistisch ausgewertet werden sollen.

Im FactGrid ermuntern wir zur Nutzung der Plattform als heuristischem Forschungstool.

  • Legen Sie hier Datenbankobjekte an, die Sie weit später erst auswerten wollen, egal welche Relevanz diese in einer Enzyklopädie oder in Bibliothekskatalogen gewinnen können.
  • Riskieren Sie provisorische Datierungen – als Arbeitshypothesen zusammen mit Ihren persönlichen Annahmegründen.
  • Wagen Sie im FactGrid unkonventionelle Aussagen, die erst einmal nur in Ihrem Forschungsprojekt interessant sind – die Software erlaubt es Ihnen.
  • Bauen Sie eigene Datenbankobjekte zu Ihren Forschungsprojekten und verknüpfen Sie diese mit allen Datenbankobjekten, an denen Sie arbeiteten, um so Ihrem Geldgebern die eigene Forschung im Paket der Datenbeziehungen vorlegen zu können.
  • Riskieren Sie im FactGrid Thesen und führen Sie diese in der Datenbank als “Mikro-Publikationen” mit eigenen Datenbank Objekt-Nummern.

Das FactGrid kommt kostenfrei – wie geht das?

Die Software ist frei verfügbar und befindet sich in einer von großen Communities getriebener Entwicklung im Wikimedia-Bereich und in Zukunft in hinzukommenden Institutionen wie denen der europäischen Nationalbibliotheken, die soeben eigene Wikibase Instanzen und Tools bauen.

Das FactGrid selbst läuft unter Ägide des Forschungszentrums Gotha auf einem virtuellen Server der Universität Erfurt. Für die deutsche URL fallen jährlich €36 Domaingebühren an, die vom Forschungszentrum Gotha getragen werden.

Nutzern stehen alle Tools aus dem Wikidata Projekt zur Verfügung. Mit ihnen sind die Standardanwendungen gängiger Humanities Projekten erst einmal abgedeckt.

Da die Software open source ist, können Projekte eigene Software-Etats gezielt in Tools und Präsentationen investieren, um die es ihnen in der eigenen Forschung geht. Arbeiten Sie gerne mit einer favorisierten Softwareschmiede zusammen, so wird diese, was den Quellcode der laufenden Software anbetrifft, vor offenen Türen stehen. Bieten Sie Ihre eigenen Entwicklungen offen zur Weiterentwicklung an, und Sie können darauf vertrauen, dass Visualisierungen auch nach Ihrer Projektförderung noch fortentwickelt werden. Streben Sie dagegen eigene Softwarelösungen mit dem Ziel einer kommerziellen Vermarktung an, so bindet Ihnen die Software-Lizenz hier nicht die Hände: Sie erhalten die bestehenden Softwarelösungen frei und können eigene darauf aufbauenden Tools jederzeit kommerziell verwerten.

Was mache ich mit unorthodoxen Forschungsanliegen?

Wikibase ist bahnbrechend offen in den mit dieser Software möglich werdenden Datenbankanwendungen. Letztlich werden hier nur Beziehungen zwischen Q-Nummern hergestellt (respektive Beziehungen zwischen Q-Nummern und Daten, Q-Nummern und Raumkoordinaten, Q-Nummern und Mediendateien, Q-Nummern und Internetadressen).

Was das für Beziehungen sind, das ist für die Software irrelevant. Q1 – P1 – Q2 ist ein “Triple” und kann bedeuten “Johann Sebastian Bach (Q1) ist der Vater von (P1) Carl Philipp Emanuel Bach (Q2)”; es kann genauso gut bedeuten “Der Brief des Archivs X mit der Signatur YZ (Q1) notiert als Absendeort (P1) München (Q2)”

Q-Nummern können für alles nur Denkbare vergeben werden – Personen, Dokumente, Ereignisse, Ideen… Sie selbst legen fest, was für P-Nummern Sie definieren wollen, um Aussagen Ihres Interesses zu Ihren Datenbankobjekten zu treffen. Im System kristallisiert sich letztlich erst mit den Aussagen, die Sie zu Ihren Objekten treffen, heraus, welcher Art Objekte das sind. Das heißt: Sie benötigen kein Kategoriensystem, das Ihnen im Vorhinein klar sein muss, um mit der Datenbank zu arbeiten. Sie treffen Aussagen dann, wenn Sie sie plötzlich setzen wollen, und sehen zu, wie sich diese zu einer kritischen auswertbaren Masse akkumulieren.

Alle Aussagen lassen sich “qualifizieren” – “Johann Sebastian Bach (Q1) war verheiratet mit (P2) Maria Barbara Bach (Q2) mit Beginn am (P2) 7. Oktober 1707 (Datum) bis zum (P3) ca. 5. Juli 1720 (Datum).” Alle diese Aussagen lassen sich wiederum beliebig mit Quellenaussagen versehen – “das geht hervor aus (P4) dem Kirchenbuch von… (Q3)”, “das wird so behauptet in (P5) der Bach-Biographie XYZ (Q4)”.

Das System lässt jederzeit konkurrierende Behauptungen zu. Sie werden einfach mit ihren unterschiedlichen Quellen eingebracht und können dabei beliebig untereinander bewertet werden. In ihrer weiteren Nutzung im System gewinnen sie ihre eigentliche Bedeutung.

Letztlich lassen sich auf damit beliebige normalsprachliche Aussagen generieren; vor allem aber öffnet dies die Tür in die Welt global handhabbarer Aussagen. Für das System spielen sich alle Aussagen nur als Verknüpfungen von Q-Nummern mittels P-Nummern ab. Sie selbst belegen die Q- und P-Nummern mit “Labeln” in den Sprachen, in denen Sie kommunizieren (Datums- und Mengenangaben verwaltet das System bereits in beliebigen globalen Standards mit automatischer Umrechnung in jede Richtung); so das Geheimnis, das es möglich macht, dass auf Wikibase Plattformen Autoren in ihrer jeweiligen Sprache eingeben und andere Nutzer diese Informationen in beliebigen Sprachen auslesen.

Welche Tools stellt die Software mir zur Verfügung

Datenbank-Eingaben können sukzessive erfolgen: Sie wollen über ein Objekt eine neue Aussage treffen? Rufen Sie das Objekt auf, gehen Sie ans Ende der Eingabe-Seite, wo “Aussage hinzufügen” steht, und beginnen Sie die Aussage in Ihrer Lieblingssprache – das System ergänzt mit Auto-Complete die gesuchte Aussage und sucht sich die P-Nummer dieser Aussage heraus. Tippen Sie in das sich nun eröffnende Feld ein, auf welches andere Objekt diese Aussage laufen soll oder auf welches Datum (in Ihrer Sprache) und die Software macht ihnen immer präzisere Vorschläge, noch während Sie tippen.

Datenbank-Eingaben können zweitens massenweise und automatisiert in gängigen Formaten aus Excel-Listen oder CSV-Listen erfolgen. (Dies ist die Eingabemaske und dies die kurze Anleitung dazu.)

Datenbankabfragen geschehen im Gegenzug mit “SPARQL“, eine Such-Sprache, die (leider) nicht einfach zu bedienen ist, die es jedoch nun erlaubt, die Datenbank beliebig komplex zu befragen – komplexer als jede Standard-Suchmaske Ihnen das erlauben würde. SPARQL-Nutzer müssen nicht unbedingt den SPARQL-Quellcode schreiben können. Man bedient sich in der Regel einer passenden Musterabfrage, die man in einer Sample-Queries-Liste findet, und tauscht hier die Suchbegriffe aus.

Weiß man exakt, welcher Art Suchanfragen Benutzer des eigenen Projektes durchführen können, so lassen sich Suchschablonen wie in herkömmlichen Katalogen bauen, die die Anfragen an das System unterhalb der Benutzeroberfläche in SPARQL formulieren.

Im Software-Paket finden sich Darstellungen auf Landkarten, Timelines, in Netzwerken, genealogischer Beziehungen, Diagramm-Darstellungen und so fort. Für diese Anwendungen müssen keine Softwarepakete eigens installiert werden. Sie formulieren in Ihrer SPARQL Anfrage, was für eine Darstellung Sie wünschen.

Einen sehr schönen Überblick über Visualisierungen gibt das Scholia Projekt auf der Wikidata Plattform.

Was mache ich, wenn ich ganz eigene Datendarstellungen auf meiner eigenen Plattform realisieren will?

Das sollte technisch kein Problem sein. Uwe Jung demonstrierte, wie eine Benutzeroberfläche der FH Potsdam Wikidata als Datenrepositorium benutzt, ohne dass die Benutzer mitbekommen, auf welche Repositorien die Abfragen zugreifen.

Nichts spricht dagegen, das FactGrid als ein externes Repositorium zu benutzen und das eigene Forschungsprojekt auf dem Server der Heimat-Uni aufzubauen und dort die Benutzer mit gezielten Datenbankzugriffen unter einer typischen Suchschablone zu bedienen.

Das FactGrid lizenziert Daten grundsätzlich CC0 – heißt das nicht, dass ich alle Rechte an meiner Forschung aufgebe?

Positiv gesprochen bedeutet die Entscheidung für die Creative-Commons-0-Lizenz, dass Sie die vollen Rechte an allen (Ihren) Daten behalten. Vor allem aber bedeutet die CC0-Lizenz, dass Ihre Daten frei nutzbar werden und damit weniger Gefahr laufen, als Forschung obsolet zu werden.

Einige grundsätzliche Erwägungen dazu: CC BY 4.0 ist auf den ersten Blick die Lizenz, die Wissenschaftlern näher liegt: Man gestattet die freie weitere Nutzung, wenn man dafür fachgerecht zitiert wird. In der Praxis lässt sich diese Lizenz bei Textveröffentlichungen (wie diesem Blogbeitrag) noch gut durchsetzen. Hier ist es klar, wie man den Text, den man verfasste, und Thesen, die man bahnbrechend formulierte, zitiert sehen möchte: Mit einem Verweis auf den eigenen Namen, mit dem Titel der Publikation, dem Publikationsort und dem Datum. Wie aber soll das Datenzitat (etwa die Information, dass ein Brief im Juni 1753 von Paris nach Berlin ging) innerhalb einer Visualisierung erfolgen, auf Strichen, die auf einer Landkarte erscheinen? Wie sollen Benutzer Datensätze zitieren, wenn diese nur zum Teil Informationen aus Ihrem Forschungsprojekt enthalten, zum größeren Teil jedoch andere Datenbankinformationen etwa aus einer öffentlichen Archivdatenbank bieten, die Sie einspielten? Noch größere Probleme stellen sich bei Daten mit Lizenzen, die eine “share-alike” Klausel bergen: “Diese Daten sind frei verfügbar, wenn die Nachnutzer sie genauso frei verfügbar halten.” Das klingt erst einmal nach nachdrücklicher freier Nutzung. Wie aber soll ein Nachnutzer sicherstellen, dass wiederum seine Nachnutzer sich an die von Ihnen gewünschte Lizenzvereinbarung halten (insbesondere, wenn dieser Nachnutzer auf CC0 geht)? Nachnutzer sind gut beraten, keine Daten aus CC-BY oder CC Share-Alike Lizenzen zu verwenden.

In der Praxis der größeren Kooperationen mit Wikidata und der Deutschen Nationalbibliothek kann es darum nur eine Option geben: Genauso frei Daten zur Verfügung zu stellen, wie die Kooperationspartner diese tun: CC0, das heißt ohne, dass sichergestellt wird, dass Nachnutzer noch exakt angeben, wer die Daten einzeln erhob, und was Drittnutzer wiederum mit diesen Daten tun dürfen.

 
Die maximal offene Lizenz heißt in der Praxis durchaus nicht, dass FactGrid Daten Daten ohne Autorschaft sind, ganz im Gegenteil. Wir legen es nahe, Forschung als Forschung zitierbar zu machen und gehen davon aus, dass Wikidata und die GND Zitiervoschläge gerne übernehmen. Alle Datensatzänderungen werden von der Software für alle Nutzer sichtbar an bei uns Real-namentlich gebunden. Haben Forschungsprojekte substanzieller an einem Datensatz gearbeitet, vermerken Sie dies jederzeit mit einer eigenen Notiz Ihrer Arbeit im Datensatz. Jeder kann die Datenbank danach befragen, welche Datensätze im Rahmen eines bestimmten Projektes bearbeitet wurden.

Tatsächlich ist es für Datenbanken wie Wikidata oder die Gemeinsame Normdatenbank der DNB, die GND, ungemein interessant, Daten aus dem FactGrid als dort erstmals publizierte zitieren zu können – es steigert die eigene Datensolidität, wenn sich Forschung benennen lässt und es gibt beiden Institutionen eine Plattform auf der möglich wird, was auf der eigenen Plattform ausgeschlossen bleibt.

Was geschieht, wenn ich mit meinen Daten auf einer anderen Plattform weiterarbeiten will?

Da Sie Ihre Daten ohne Copyright-Aufgabe einpflegten, können Sie sie in beliebig vielen anderen Projekten parallel laufen lassen. Wir sind hier gerne “nur” ein “Inkubator” für Forschungsdaten.

Was passiert, wenn es zwischen FactGrid-Benutzern zum Streit über ein “korrektes” Datum kommt?

Die Software lässt es zu, einander widersprechende Daten zu handhaben – das ist gerade im Feld der Geschichtswissenschaften interessant, wo wir oft dokumentarisch belegte divergierende Informationen haben, ohne noch ermessen zu können, welches die korrekte Aussage ist. Namen werden in verschiedenen Schreibungen gehandhabt; Geschichtsschreiber widersprechen einander.

Die Software erlaubt es, die widersprüchliche Lage abzubilden; sie erlaubt es, Objekte mit Dutzenden Namensschreibweisen zu belegen – es sind nur verschiedene Schreibweisen zu ein und derselben Q-Nummer.

Divergierende Aussagen lassen sich unterschiedlich einstufen – etwa als das derzeit autoritative Datum gegenüber Varianten, für die allein die verschiedenen einander widersprechenden Quellen sprechen.

Sollten zwei Forscher zu unterschiedlichen Befunden kommen, so ist das letztlich der Fall, auf den es jedes Projekt anlegen sollte. Das viel größere Risiko ist die hypothetische Entscheidung, die Sie in Ihrem Projekt treffen, während ein Projekt auf einer anderen Plattform das Rätsel löst und mit dem belastbaren Datum weiterarbeitet, ohne dass Sie davon auch nur erfahren, geschweige denn die Korrektur Jahre nach Ihrer Projektlaufzeit noch einpflegen können.

Warum sollte ich in meinem Projekt Transparenz schon von Anfang an riskieren?

Das dürfte die härteste Frage sein, die Projekte im Moment davon abhält, sich der eröffneten Ressource bereits zu bedienen. Die Alternative ist die Ressource, die bis zum Publikationstermin gegen Projektende nur unter Passwörtern dem Team zugänglich ist. Kein Konkurrenzprojekt kann Ihnen in der verdeckten Arbeit an Ihrer Plattform Befunde wegschnappen, so die Theorie. Auch sieht niemand, wo Sie erst einmal Fehler machten, die Sie erst im Verlauf korrigierten. Niemand erfasst, was Hilfskräfte eingaben und was dagegen Sie als Projektleiter verantworten – so die Vorteile des intransparenten Arbeitens auf einer Plattform, die erst am Projektende online geht und die keinen Blick in die Versionsgeschichten der Datensätze zulässt geschweige denn in die tägliche Projektarbeit.

Die transparente Forschung, zu der das FactGrid einlädt, bietet indes ganz eigene Sicherheiten: Wenn Sie bahnbrechend ein bestimmtes Dokument auffinden und einen bestimmten Zusammenhang herstellen, dann ist der Editiervorgang, mit dem Sie den Datensatz mit Aussagen bestücken, Aussage um Aussage datiert und jeweils mit Ihrem Namen verbunden. Macht morgen jemand im selben Archiv dieselbe Entdeckung, werden Sie im FactGrid nachweisbar und fälschungssicher notiert haben, dass Sie den Zusammenhang schon einen Tag zuvor öffentlich zugänglich machten.

Die kollektive Plattform spricht im selben Moment die Einladung zur Kooperation aus. Machen Sie anderen Teams klar, woran Sie arbeiten und erlauben Sie noch auf Ihrer Plattform die Kontaktaufnahme mit Ihnen!

Die Risiken des vermeintlich sicheren und erst am Ende der Projektzeit online gehenden Internetauftritts sind gravierend: Die Zeit für einen Austausch mit den Nutzern ist mit der finalen Publikation abgelaufen. Der Internetauftritt erfolgt in den letzten Wochen Ihres Projektes, währen alle anderen letzten Arbeiten laufen und für konzeptionelle Änderungen kein Raum mehr bleibt. Ist überhaupt kein Digital Humanities-Part in der Forschung einkalkuliert, so bleibt unklar, was mit den Daten noch geschehen soll, die Mitspieler in Word-Dateien, Excel-Listen oder eigenen Softwarelösungen ganz für sich sammelten – sie noch irgendwo einzupflegen, hat niemand mehr die Kraft. Man kann nur noch hoffen, dass Leser der abschließenden Buchpublikation in dieser alle Fußnoten auf Korrekturen hin durchkämmen, die im öffentlichen Informationsstand nötig werden und diese dann in allen Bibliothekskatalogen und in den verschiedenen Wikipedia-Projekten durchführen. Das Risiko sind hier Bücher, die an der kollektiven Datengrundlage vorbeigehen und DH-Projekte, die nach ihrer Publikation mangels weiterer Pflege in wenigen Monaten veralten.

Die Zukunft sollte darin liegen, dass wir die Datenlage, derer wir uns selbst in der Forschung bedienen, eigenverantwortlich bearbeiten können. Um dies wiederum ohne Gefahr der Aufgabe wichtiger Befunde tun zu können, müssen wir es Forschern erlauben, ihre Forschungsleistung transparent und zitierbar sichtbar zu machen. Hierfür ist Wikibase besser als jeder andere Software ausgerüstet.

Wie nutze ich das FactGrid konkret?

Das FactGrid hat keine unsichtbare Tiefenschicht. Jeder kann die Datenbank befragen und die Abfragen bieten dabei eingeloggten Benutzern dieselben Antworten wie fremden. Alle Datensätze sind in allen Versionsstufen öffentlich präsent. Der eigene Benutzeraccount bietet beim puren Auslesen von Daten nur den Vorteil, dass Nutzer nun die Software auf die eigene Sprache umstellen können.

Wer eigene Daten einspeisen und mit ihnen auf der Plattform arbeiten will, benötigt dagegen einen Account. Diese werden unter Klarnamen von den Administratoren vergeben. Die Software bietet ein Link zur Account-Anforderung. Wir vergeben ansonsten Accounts auf (Email) Anfrage. Projektleiter erhalten administrative Zugänge, mit denen sie selbst Benutzerkonten vergeben können.

Einmal eingeloggt kann man sowohl Daten in Massen eingeben wie an beliebiger Stelle Datenkorrekturen vornehmen. Alle Eingaben werden an den Benutzer-Account gebunden, der sie tätigt, und können von jedem anderen zurückgenommen werden, was wiederum als exakt diese Rücknahme namentlich dokumentiert und mit einem Zeitstempel versehen wird – für eingeloggte und nicht eingeloggte Benutzer gleichermaßen sichtbar.

Wer an einem komplexeren Projekt arbeiten will —

  • das kann die persönliche Familienforschung sein,
  • das kann eine einzelne Visualisierung im Rahmen einer Seminararbeit sein,
  • das kann genauso gut eine Eingabe von Tausenden von Datensätzen in einem auf mehrere Jahre laufenden Forschungsprojekt mit mehreren Mitarbeitern sein

— ist gut beraten, noch bei der Account-Eröffnung den Austausch mit der Plattform zu suchen. Eine Kooperationsvereinbarung kann dem eigenen Projekt Rückhalt gegenüber Geldgebern geben; doch kann auch einfach nur ein Blogbeitrag mit einer Projekteröffnung interessant sein. In jedem Fall sollten andere auf der Plattform von Ihnen wissen. Spannend wird die kollektiven Datenbank, wo man die Vorarbeiten anderer nutzt und Mitspieler aus anderen Projekten anregt, gute Modelle zu übernehmen. Nötig ist die Abstimmung nicht; praktisch aber trägt sie zur Verbreitung der eigenen Arbeit bei – zu Suchanfragen, die man selbst so schnell nicht zu formulieren weiß, zu Visualisierungen, an die man nicht dachte, zu Hilfe, wo man sie benötigt.

Die Software verspricht eine neugierige Nutzung im breiten Netz und dieses Netz sollte man mit Forschung ansprechen.

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).
Jack Kirby, “The Fourth Dimension is a many splattered thing!” from Alarming Tales, 1 (September 1957).

Memorandum of Understanding between the University of Erfurt and the German National Library – to base the FactGrid on GND data in a joint project

German Version

We are proud to announce a new and massive Wikibase project that should keep a large community busy for far more than a year: Last month the president of the University of Erfurt, Prof. Dr. Walter Bauer-Wabnegg, and Dr. Elisabeth Niggemann, director-general of the German National Library in Frankfurt and Leipzig (DNB) signed a memorandum of understanding that aims to bring GND data into the FactGrid – on a grand scale.

The GND, the German Integrated Authority File, is an authority file of millions of persons plus corporate bodies, conferences and events, geographic information, topics and works – designed to shape the exchange between libraries, archives and academic projects in the DACH countries of Germany, Austria and Switzerland.

integrating the GND into the FactGrid had been our constant topic of discussion during the last year. A Wikibase instance becomes a cool thing to contribute to, as soon as it becomes the research tool that you would use yourself in your research. GND data links into the world of open data; they clarify who or what you are speaking of in your research in all German-language contexts – and they will reach out to the other global authority files and to the universe of library data.

In April 2018 it became clearer that the FactGrid would eventually be one of several Wikibase instances which could and should in this case aim for a larger federation. Early in June it transpired that the German National Library was on its way to test Wikibase in a software evaluation, with the aim to run possibly about ten Wikibase instances in a constant exchange with each other. That was when we contacted the DNB with our own agenda to import their data. We wanted to try, so that our proposal, could become a platform for “original research” – a platform without GND or Wikidata criteria of notability – in the evolving network of Wikibase platforms. Users will be allowed to create Q-Numbers for infants who died right after birth on FactGrid, and the GND and Wikidata will be free to decide under their criteria of notability and relevance, whether they would like to use our information – information they can now quote as original research from the FactGrid platform (with the detailed information of the projects behind this research).

Whilst the GND is CCO and free to be copied, the open joint venture with the German National Library aims to bring transparency into the data input. The more transparency we can bring into all the design decisions in this early stage, the better the wikibase platforms we are heading towards, will eventually be able to communicate with each other.

Now a team has to be formed. The German National Library and the Gotha research institutions of the University of Erfurt will send members into the team. The question is: Will we be able to broaden this team? We should have experts from the Wikimedia communities on board – people who know Wikibase and Wikidata, people who are used to community work on a regular wiki.

  • We would like to attract people who know how to formulate SPARQL searches and who will be able to test data models and make suggestions for the improved data models we should use, in order to handle the massive data sets we are expecting.
  • We’re looking for Wikibase experts who know how to bring in tens of millions of records into a Wikibase installation, and who know how to interconnect these records with genealogical and geographical links.
  • We do not yet know how we will keep the FactGrid manageable with respect to the wave of doublets and name parallels we are facing: The GND has these name parallels in unprecedented numbers. We will have to find ways to quickly inform researchers whether a person they have found in a document is already on the FactGrid or whether they will have to create the item. The hunt for items to be merged will become a permanent issue and we do not yet know how to technically support a community on this collective quest.
  • We will create new and complex fields of expertise: Millions of personal data sets will come with career statements. The FactGrid will turn all these statements into Q-Items, which we will have to organise in order to allow sociological searches for instance. The FactGrid project on historical jobs and their evolution will be only one of these projects.
  • We need players with Wikipedia experience: Though we will restrict ourselves to clear name accounts, we widely invite users with professional to private ambition to join the platform with their projects – whether they are focused on private genealogy or on publicly funded historical research.
  • We will have to provide a simplified FactGrid user interface that will bypass the SPARQL QueryService and the mushrooming Wikibase input pages. Magnus Manske’s Reasonator might become our standard interface for regular users, who will access the FactGrid as if they are accessing library catalogues – through organsied input forms.
  • We will eventually need help with database maintenance. It is particularly unfortunate that our project is primarily the work of historians, who do not always have a keen eye on how to optimally supply this technology.

The FactGrid will grow – and it will offer plenty of space for people to develop their own projects within this growth.


Scan of the Memorandum of Understanding (in German)


More

  • Barbara Fischer & Jens Ohlig, “Neues Testfeld für Wikibase: Eine Bundesbehörde geht auf Expedition im Wikiversum.” 2019-05-09 at https://blog.wikimedia.de