Beam-me-up: A Tool for Importing Wikidata Entities to FactGrid

The WikibaseMigrator, also known as the Beam-me-up tool, automates the complex process of transferring data from Wikidata to FactGrid. Traditional methods such as manual creation or imports through QuickStatements and OpenRefine often prove to be time-consuming and error-prone. This new tool simplifies the entire process by automatically mapping properties and items between the two Wikibase instances, having already performed over 18,000 successful edits.

How to Migrate Entities

Starting the migration requires to select the entities to migrate. Here the input of the single ID or a list of ID is possible

The migration process requires only Wikidata IDs as input and offers three flexible input methods: you can enter a single ID, provide a comma-separated list of IDs (make sure you do not end with a comma), or use a SPARQL query. After entering the IDs in the input field, a preview table displays the selected entities with their English labels, allowing you to verify your selection before proceeding.

When you click “Run matching!”, the tool queries the selected entities’ data from Wikidata and begins the translation process. It extracts all properties and items from the dataset and searches for corresponding mappings in FactGrid. Using these mappings, the tool translates the entities to FactGrid entities. If an entity already exists in FactGrid, the tool will augment it with the new statements from Wikidata. This process may take several seconds, depending on the number of entities selected and the complexity of their statements.

Here the translation can be checked before starting the migration. Additionally a summary or the project ID can be provided.

Once the translation is complete, you will see the results for each entity, indicating whether the process will create a new entry or augment existing ones. At this stage, you can add a research project ID (P131) if the entities belong to your research project, which adds a corresponding statement to each migrated entity. As with any wiki edit, you can provide a summary explaining the reason for the migration—this is particularly recommended for large imports.

The Migration Process

After reviewing the translation and clicking Beam me up!, the actual migration begins. A progress bar keeps you informed of the process. Upon completion, you receive a comprehensive overview table of the created or augmented entities, which you can download for your records or further additions of new statements you want to make. This table includes both the Wikidata and FactGrid IDs, along with detailed migration information.

The tool handles property type mismatches between Wikidata and FactGrid through automatic type casting where possible. For instance, it can convert string values to quantities and manage monolingual text to string conversions and vice versa. Any such transformations are documented in the migration details column of the results table.

The migration result is shown as table with the IDs of Wikidata and FactGrid.

If a migration fails, the tool provides a separate table showing the affected Wikidata ID and the reason for the failure. During testing, most failures were related to existing sitelinks, as Wikidata supports defining sitelinks to redirects, which FactGrid does not.

 

Entity Augmentation and Advanced Features

The tool employs a sophisticated approach to augmenting existing entities, optimized for multiple augmentation cycles. Statements are considered equal if their main values match and either has no qualifiers or their qualifier sets are identical or one qualifier set is empty. In such cases, the tool merges references and qualifiers intelligently, preventing duplicate statements that might occur with the default merging strategy.

This feature proves particularly valuable when migrating interconnected data, such as family relationships. For example, when migrating a group of related persons with Mother (P142), Father (P141) and Child (P150) relationships, the tool can handle the circular dependencies effectively through multiple migration passes.

For example when migrating the list of Q81642270, Q81642507, Q28085 (assuming that they are not already in FactGrid) the statements with property mother (P25) and father (P22) will not be migrated as the target entities do not exist at the time of translation and thus do not have a FactGrid ID yet. But applying the migration a second time to the same list of entities leads to the augmentation of the now existing FactGrid entities allowing to migrate the statements with the properties mother (P25) and father (P22).

 

Technical Implementation and Additional Tools

All migrations are performed under the user’s account credentials, with each edit tagged to indicate it was made using the tool.

WikibaesMigrator Edit Log Example

The Wikidata ID of the original entity is automatically added as a sitelink to newly created entities, ensuring proper linking between the two databases and facilitating future augmentations. It should be noted that the back reference is also configurable and can also be configured as external id with a property which would be a cleaner solution.

To complement the WikibaseMigrator, a Wikidata bot called FactGridSync periodically queries FactGrid’s latest edits and updates the corresponding Wikidata entries with FactGrid IDs. This synchronization covers all FactGrid entities, not just those migrated using the Beam-me-up tool.

 

Example of the FactGridLinker annotations
FactGridLinker adds the FactGrid ID to each Wikidata entity page or a link to the Beam-me-up tool if the entity does not exist

For those interested in enhanced Wikidata functionality, I implemented the UI extension FactGridLinker, initially for debugging, that simplifies checking whether Wikidata entities exist in FactGrid. This user script can be enabled through your commons.js configuration (see here for details).

The WikibaseMigrator’s versatility extends beyond the Wikidata→FactGrid

relationship—it can be configured to work between any Wikibase instances with proper configuration.

It was successfully used for subsetting the CEUR-WS data into its own wikibase instance with just a few queries that defined entities of the subset and was performed in under two hours.

Should you encounter any issues while using the tool, you can report them on the my talk page or by opening a issue on GitHub.


Image Maniere universelle de M. Desargues, pour pratiquer la perspective par petit-pied… Planche 4 Gallica France

At least a make shift solution: The “Julian calendar stabiliser”

My last blog post triggered a couple of responses on Twitter. It seems I touched a problem that will not be solved that easily.

Save dates as Julian on your Wikibase (manually or, with the /J switch, in your QuickStatements mass input) and your Wikibase will be able to handle these dates correctly in any mixed bag of Julian and Gregorian dates. It is nice that the Query Service is able to produce straight timelines out of any such mixed bag, but immensely problematic that you will be quite unable to get the original Julian dates back in regular Query Service downloads. Blazegraph, the tool that is working behind the Query Service, does its job on normalisations of dates, and these are, of course, performed in the superior Gregorian calendar. Wikibase Query Services are hence on their way to produce loads of unprecedented arithmetical Gregorian dates in environments that have been solely Julian so far. We will first be puzzled by dates that strangely differ from those we fed into these machines — we will have to understand that they have silently added days on them to reach their Gregorian equivalents. Handle these artefacts as correct Gregorian dates, though they are without evidence in the historical records — do not feed them as Julian into any Wikibase because that will immediately expose them to the next round of Julian to Gregorian conversions wherever a Query Service will spot them.

SPARQL queries can actually produce the complexity of the Wikibase they are accessing, but that requires quite some scripting skills. Tagishsimon gave the following script that helped him to get well informed dates from Wikidata in this Twitter response:

Bruno Belhoste applied this script in the following FactGrid query, which will be extremely useful in all future searches on our database. The table gives you the birthdays of members of the French Academy — a typical “mixed bag” of dates that shows all imaginable challenges of different calendars and the various precision statements:

Change the parameters and you will get the dates you are interested in with all the information you will need to process a mixed bag of historical dates from the Wikibase of your choice.

A make shift solution: The “Julian Calendar Stabiliser”

We agreed that we have to stabilise Julian dates on FactGrid under these conditions. All Julian dates will be translated to Gregorian sooner or later on our Query Service. Users must, hence, remain able to get the original Julian information side by side with their (secretly Gregorianised) searches. The simple solution is a string repetition of the Julian statement you want to make. The Query Service does not touch strings, chains of characters and numbers; it will give you the original Julian statement which you can use in other contexts as the very dates you saw in your documents:

Johann Sebastian Bach’s birthday with the “Julian date stabiliser” (see it in the data set)

This is not the ideal solution. One would rather like to have a calendar sensitive Query Service that produces dates as stated on your Wikibase; but it is at least a pragmatic stabilisation to keep Julian dates intact in the waves of transformations and deformations which we are likely to witness in the new world of data processing.

Are our Wikibase QueryServices about to mess up two millennia of historical dates?

It was in February 2019 at a conference dinner of medievalists in Jena when I was first confronted with the calendar problem which Wikibase had been posing ever since it had digested its first Julian calendar dates. I had given a Wikibase demonstration earlier that day and now I was sitting next to a medievalist who was ready to destroy me: “Wikibase”, he stated, “is a genuine disaster without anyone understanding it.”

I demanded to hear why that should be the case and the man asked me to show him just one medieval date from Wikidata. I had activated my phone and landed on biography c. 1500.

“See that small print?” he asked, “these dates are all noted as Gregorian before 1584.”

The qualifier was indeed peculiar. Why would they set a Gregorian date before 1582 and then mark it as such? “Well, you know, that the Gregorian calendar was only introduced in 1582, do you?!”

Of course I knew. I am an 18th-century person and Britain had introduced this calendar as late as 1752. The reform had by that time to close a gap of 11 days. But I could also point out that Wikibase allowed the fast correction: “You can easily switch between the calendars, and the machine will actually understand the implications on any timeline” I showed him my screen:

The man was in agony: “Too late. Wikidata is already in big shit”. I realised that I was lacking the full astronomical background and that I did not know the story of these peculiar Wikidata redactions.

Why we needed the Gregorian calendar in the first place

Both, the Julian calendar of 46 BC and the superior Gregorian calendar first introduced in 1582, are approximations. A solar year is one circle around the sun whilst the globe is spinning at about 365.2422 revolutions per year — year after year our planet ends its tour with a different slice pointing towards the sun. We are, in fact slowing down, thanks to the friction which the moon’s gravitation is generating in a constant movement of ebbs and tides, but that is another story. 365.2422 turns per year is our present spin more or less exactly but difficult to generate in a procedural long term pattern of constant adaptations.

The Julian calendar, as it was introduced under Julius Caesar in 46 BC, added one day every four years — in the so called leap years — a rule that boiled down to an additional quarter of a day per year. The approximation of 0.25 days against 0.2422 missed its mark just by 0.0078 days per year, less than a hundredth of a day — negligible one might think — but that one hundredth of a day is a day in a hundred years. In a millennium this discrepancy is piling up to 7.8 days, in two millennia to half a month, moving Christmas further and further away from the longest night until we can finally celebrate Christmas and Easter on the same day; and that was why the Gregorian calendar was finally introduced in 1582 with its far more complex regime of leap years:

  • add one day every four years (as you did under the Julian calendar to create a year of 365.25 days)
  • omit every leap year that is divisible by 100 to get a lower number
  • let this leap year, however, happen if it is divisibly by 400 in order to get a year of 365.2425 days.

The Gregorian calendar reduced the aberration to a surplus of 0.0003 days per year — it will now take 3333 years until we need an additional day to be back in tune with the solar year. The Vatican in Rome adopted the calendar on the 4th of October 1582 — jumping over night into Friday the 15th of that year. Christianity, however, was at that point no longer ready to obey a Papal decree. Eastern Orthodox churches stayed on the Julian calendar right into the the 20th century; Protestant territories and realms would decide one by one. Prussia (with its complex ties into catholic Poland adopted the new calendar in 1612 whilst most of the other Protestant territories stayed Julian for the next 88 years. The United Kingdom took the step in 1752. Lithuania, Russia, and Greece were to switch as late as 1915, 1918 and 1923 respectively.

The following map is from reddit:

When Europe switched from Julian to Gregorian calendar.
byu/coneyislandimgur ineurope

…and it is far from getting the full complexity. The following list gives the growing FactGrid table:

Europe was fragmented. Travelling across Germany in 1699, you could date your letters switching back and forth at every customs house on your tour:

Map of the Holy Roman Empire 1648. Wikimedia Commons

How we solved the problem — and created an even bigger one

Wikibase is a bright software. The tools — the QueryService and QuickStatements — are (or were) not immediately that bright, and that caused the mess the medievalist had noted. QuickStatemens, the tool for mass imports, simply did not offer a Julian calendar switch before February 2023. Instead it would mark all dates as Gregorian without asking — which, looking backwards, was not all that bad…

…why could we all live with the erroneous labelling of Julian dates as Gregorian on Wikidata? Because Wikidata was with this negligence basically doing what we all had been doing up to that point.

Johann Sebastian Bach was born on the 21st of March 1685. Germany’s central database, the GND, is stating this date up until now without the slightest remark on the calendar. The date is Julian because Eisenach’s church register was keeping records in the Julian calendar for another 15 years. The composer himself will not have shifted his birthday to the 31st of March in 1700, the year of the great reset. We all ignore the shift and keep copying dates from documents without any interference. Calendar experts might be interested in the “real” day and they can create Julian/Gregorian calendar matches in those rare cases in which they have to create an exact timeline of events with dates of both calendars.

The Wikidata community had been unaware of the problem. The Gregorian label on all the Julian days was foolish, but the input was actually stabilising our historical tradition as the QueryService will not do anything odd with dates that are entered as Gregorian.

I was far from seeing these advantages after my conversation of 2019 and that was why I warned the PhiloBiblon team in 2022 that QuickStatements would label all their Julian dates as Gregorian against all better intentions once they were imported to FactGrid. Charles Faulhaber immediately asked their programer, Josep Maria Formentí, whether he could not take a look into QuickStatements to solve that little problem. Weeks later Josep introduced the /J-switch that is now available to mark any date as Julian in mass inputs:

+ 1751-06-16T00:00:00Z/11/J

You can now feed thousands of medieval or early-18th-century British dates into your Wikibase and your machine will present all these dates in timelines in perfect synchrony with Gregorian dates. This is extremely nice if you are editing a correspondence whose partners were signing their letters under various calendars. Your machine will give you the exchange of letters in their true course.

So why the alarm?

Wikibase is an intelligent software; it brings objectivity into your statements. Feed a Julian date into your Wikibase and that day will be noted as Julian on the Wikibase itself.

Things get messy wherever we retrieve Julian dates from the QueryService, since this is where the production of funny (and eventually of erroneous) dates will be begin. The QueryService will convert all Julian dates into mathematically correct Gregorian dates.

Martin Luther is known to have died on the 18th of February 1546 — under the Julian calendar, that needs not to be stated, and our Wikibase is giving that date without any calendar stamp on it. But ask the QueryService for Luther’s birthday and it will tell you that the church reformer actually died on the 28th of March, 10 days later — a Gregorian calendar date (without indication) (no big issue you might think, now that you know).

And now think of masses of data which we will be moving between Wikibases in the brave new world of “federates Wikibases”. If there are “Julian” dates among them, then these will get secret additional days wherever they are extracted with the help of a regular SPARQL-Query on the QueryService.

This is what will happen to Luther’s date of death as it is now no longer a subject of safe copying. We will see it in an increasing number of variants — namely as:

  • 18 February 1546 (Greg.) — mistaken QuickStatements input artefact
  • 18 February 1546 (Jul.) — the historically correct date
  • 28 February 1546 (Greg.) — unorthodox but correct Wikibase QueryService output
  • 28 February 1546 (Jul.) — Wikibase output mistakenly saved as Julian
  • 10 March 1546 — the previous converted to Gregorian
  • 20 March 1546 — the previous after the next im- and export

and so on and so on.

Can we stop the wave of uncontrolled additions of days on Julian calendar dates?

I am not quite sure how. We need a QueryService that will never ever offer a historical date without the corresponding calendar statement (now that we have a machine that does both calendars).

But not only the QueryService is posing a problem here. Our Wikibases should have a third option, because our documentary evidence is usually lacking calendar information. Eisenach’s church register of 1685 is using the Julian Calendar (without further notice), that is something we can determine — but we cannot say what calendar an author of a typical letter was using in 1685 if that date comes without a localisation. Our documents do not tend to have calendar statements on them.

What we need here is a third — a “calendar format unknown” — option. It’s complicated, I am afraid.

Links and more

  • Header image from Ολυμπία δώματα, or, An almanack for the year of our Lord God 1752 (London: Printed by T. Parker, for the Company of Stationers, 1752), from the digitisation at Archive.org
  • English Wikipedia List of adoption dates of the Gregorian calendar by country https://en.wikipedia.org/
  • See also: Maniphest T207705, Implement the Extended Date/Time Format Specification, https://phabricator.wikimedia.org/T207705
  • Lydia Pintscher, calendar model screwup, 30 Jun 2015. [https://lists.wikimedia.org/hyperkitty/list/wikidata@lists.wikimedia.org/thread/Y7OEHUYV66DHRVZ6JCSODWAYZ25SLUHM/ https://lists.wikimedia.org/hyperkitty]
  • Julian and Gregorian dates from Wikidata, question asked on https://opendata.stackexchange.com/, Apr 18, 2018 at 0:33 [https://opendata.stackexchange.com/questions/12723/julian-and-gregorian-dates-from-wikidata https://opendata.stackexchange.com/]

Filling a Wikibase instance with millions of data

As more and more Wikibase instances are cropping up we are seeing attempts to start them with masses of data from already existing data bases that want to switch to the new software.

Experimenting I tried to find a faster way to insert a huge amount of items into a Wikibase instance. I have not been able to insert more than two or three statements per second using the ‘official’ tools, such as QuickStatements or the WDI library.

Therefore, I am inserting the data directly into the MySQL database used by Wikibase.

The process consists of these steps:

  • generate the data for an item in JSON
  • determine the next Q number and update the JSON item data accordingly
  • insert data into the various database tables

However, if you do this without a transaction it is still terrible slow. In my setup only 120 items per minute. However, if I wrap the inserts into a transaction I was able to insert 33,000 items/minute.

Steps to run the experiment

  mysql:
    image: mariadb:10.3
    restart: unless-stopped
    ports:
      - "3306:3306"
    volumes:
  • Start the containers: docker-compose up and wait until you see lines ending like:
[main] INFO  o.w.q.r.t.change.RecentChangesPoller - Got no real changes
[main] INFO  org.wikidata.query.rdf.tool.Updater - Sleeping for 10 secs

For me it took a minute to insert 100 items without a transaction and 25 seconds to insert 10,000 items with a transaction.


first published at https://github.com/jze/wikibase-insert/

2018-06-11/12: Daten in das FactGrid füllen — praktischer Workshop der Gothaer Forschungsstelle Illuminatenforschung

Liebe Erfurter und Gothaer FactGrid Interessierte,

in den letzten Wochen arbeiteten wir im engeren Kreis an der Einrichtung einer WikiBase-Instanz (der Software hinter dem Wikidata-Projekt) auf dem Server der Uni Erfurt: https://database.factgrid.de/

Es gab einige technische Schwierigkeiten bei der Anpassung der Tools an die neue Server-Umgebung zu bewältigen, doch sind wir seit einigen Tagen soweit, dass die Grundausstattung läuft: QuickStatements steht für den massenweisen Datenimport zur Verfügung. Der Query Service läuft, so dass wir SPARQL-Abfragen von Daten hinkriegen. Das Design (Logo etc.) ist noch offen – was im Moment den Vorteil hat, dass alles wie bei Wikidata aussieht. Matti Blume arbeitet daran, die Datensätze des Illuminatenprojektes für den Projektstart in die Datenbank zu füllen.

Vom Montag den 11. auf Dienstag den 12. Juni 2018 wollen wir gemeinsam mit Matti Blume und Sandra Müllrick am Forschungszentrum Gotha einen Workshop veranstalten, auf dem er praktischen Unterricht in der Befüllung der Datenbank erteilen wird:

  • Wie verändert man einzelne Daten?
  • Wie fügt man über QuickStatements Datenmassen in die Datenbank ein? (Auf dieser Seite ein paar Tutorials: https://factgrid-tools.geschichte.uni-halle.de/blog/archives/811)
  • Wie legt man dazu Properties (Eigenschaften) für Items an?
  • Wie organisieren wir unsere Properties am besten (von dem weit größeren Wikidata-Projekt lernend)?

Projekte, die am Forschungszentrum Gotha und an der Uni-Erfurt im Feld Geschichte und den benachbarten Kulturwissenschaften Datenbankleistung suchen, sind eingeladen, hier praktische Einblicke zu gewinnen. Hilfskräfte, die uns beim Betreuen der Ressource in Zukunft zur Verfügung stehen wollen (und Interesse an den Digital Humanities für die eigene spätere Arbeit haben) betrifft dieselbe Einladung.

Wir werden am Montag einleitend die Software vorstellen und im Austausch mit den Anwesenden ausloten, zu welchen Projekten sie sich eignet. In einem zweiten Arbeitsblock wird es darum gehen, vorbereitete Daten in die Datenbank zu füllen und die Arbeit zu überprüfen. Das wird im Wesentlichen auch das Programm für Dienstag sein, nun mit dem Ziel, Unabhängigkeit bei der Arbeit mit der Datenbank zu gewinnen. Die Uhrzeit für den Beginn der Veranstaltung Montag-Vormittag/Mittag wird im Vorfeld bekannt gegeben.

Die Weitergabe dieser Post an Datenbank-Interessierte ist ausdrücklich erwünscht. Für Voranmeldungen bis zum Mittwoch, den 6. Juni wären wir dankbar, auch für persönliche Notizen zur Hardware: Die Software lässt sich vom Laptop (über Eduroam) bedienen. Wir wollen zusehen, dass wir Computerarbeitsplätze im Forschungszentrum für die Teilnehmer frei halten.

Mit den besten Grüßen,
Olaf Simons

[per e-mail distribuiert]

Zeitplan

Montag, 11. Juni, Pagenhaus des FZG, Seminarraum

(13:15 – 13:40) Olaf Simons: Begrüßung. Kurze Projektgeschichte, eingehenderer Blick auf die Datenblättern aus dem Illuminatenprojekt, die das erste Unterrichtsmaterial geben werden.

(13:45 – 14:15) Sandra Müllrick: Eine kurze Vorstellung von Wikidata und der Wikibase Software. Die Datenbank, die Benutzer editieren können. Versionsgeschichte, Transparenz aller Editiervorgänge über Recent Changes. Triples als Statements. Abfragen über SPARQL, mehrsprachige Datenblätter im Reasonator. Was kann eine Datenbank, was ein reguläres Wiki (wie wir es im Illuminatenprojekt hatten) nicht kann?

(14:25 – 14:45) Praxisteil 1: Wie geht man strategisch vor, wenn man Tabellen aus einem Forschungsprojekt vor sich hat und diese in eine Wikibase Datenbank überführen will? Wie legt man Properties an? Wie behält man Überblick über schon existierende Properties? Wie geht man strategisch vor, wenn man einen solchen Datenberg einarbeiten will?

(15:00 bis zum Erschöpfungsbeginn) Praxisteil 2: Koordinierter Versuch, möglichst viel unserer Daten über QuickStatements in das System zu bringen. Wenn man Tabellenspalten in massenweise Triples überführt – wie legt man die Items an, wie die Properties? (Tutorials unter anderem hier: https://factgrid-tools.geschichte.uni-halle.de/blog/archives/811). Aufgaben für Einzelne oder Zweiergruppen.

Gemeinsames Abendessen

Dienstag, 12. Juni, Pagenhaus des FZG, Seminarraum, respektive Hauptgebäude, Besprechungsraum

(9:00 – 10:30, FZG Hauptgebäude, Besprechungsraum) Planungstreffen: Sandra Müllrick Mitglieder des Forschungszentrums und der Forschungsbibliothek. Projektpläne. Welche Entwicklungen können wir aus eigenen Mitteln finanzieren – wo brauchen wir (etwa bei der Suche von Werkvertragsnehmern) Hilfe von Wikimedia, respektive der Community? Welche Entwicklungen würde Wikimedia gerne anstoßen?

(Parallell 9:00 – 11:00) Praxisteil 3: Fortsetzung von angefangenen Eingaben und individuelle Beratung, insbesondere falls Mitspieler eigene Datensätze in die Datenbank bringen möchten.

(11:15 – 12:00) Praxisteil 4: Wo befinden sich die eingegebenen Daten nun? Was kann man mit ihnen bereits machen? Vielleicht finden wir einige interessante SPARQL-Abfragen.

(12:30 – 13:00) Resümee: Ideen, Desiderate, Pläne.

 

Google Spreadsheets and Illuminati Project Data

Dataset Google Spreadsheet Content What could be visualised? Web Form to build
Documents produced in the Illuminati Order Mostly Schwedenkiste, Illuminati materials from various archives and those documents that were already published in the late 1780s Network information. Caution: We are operating with fragmented and highly selective data. Describe an archival document
Illuminati members and others A list of the c. 1350 members of the order (including names belonging into the wider context), mostly research (stated in col. AJ) by Hermann Schüttler (2016). — The geographical spread of the Order on the central European map: col. AB to be matched with col. AG
— Already existing Wikidata and GND data sets linked in cols. M and N.
— Age structure of the members col. AC
— Percentage of aristocracy cols. J-L
Write a CV
The order had a complex grade system that created careers within the order, people could also get into different official positions Give information about an Illuminati career
Illuminati Events Mostly protocols of gatherings of “Minerval Churches” under Bode’s supervision. (Missing: Hermann Schüttler’s information about gatherings of the “Minerval Church” in Frankfurt/Main) Give information about a session (as type of an event)
Organisations
Publications  from: Forschungsliteratur

Bildnachweis

The (sobering) status report of Friday 13, April 2018

[A version of this was originally posted here]

[Postscript Friday 4, May 2018: SPARQL is on, we are in the middle of our first more massiv data input]

Four months have passed since the kick-off workshop shop, and the FactGrid project has run into its first unexpected problems. We are confident that we will solve the – primarily technical – issues, but one of the lessons we have learned so far is that we will need the support of a larger community in order to situate the FactGrid Project with more impact in the Wikidata-community.

What do we want to achieve? We are still trying to launch a Wikibase installation with the aim to offer a platform for original research. Data hosted on the FactGrid will be free to be used by Wikidata. Data will leave the FactGrid database, however, with the personal authorisations of research which Wikidata is not be able to generate.

Digital humanities projects interested to work on the collective FactGrid platform will sponsor software developments with their respective DH-funding. The cooperation with Wikimedia should make sure that tools sponsored by us will become part of the wider Wikibase software package. We want to prevent island solutions.

What kinds of problems have we been facing? And where do we need you?

Problem 1: The software is free but the vital tools do not work outside the Wikidata environment.

Wikibase is – relatively – easy to install, but the central tools you need in order to get data into and out of the database – QuickStatements and SPARQL – proved to be hard wired to the original Wikidata compound. Lucas Werkmeister has managed to free Quick-Statements from these ties. SPARQL remains on his agenda. We have no idea how tools that use SPARQL (in order to create visualisations for instance) will work once we have the independent SPARQL version. The software problems have blasted our entire schedule.

Problem 2: Getting the first sets of data into the FactGrid.

We have four larger spread sheets of data from the Gotha Illuminati project which we want to use in order to create an attractive show case.

  1. Google Spreadsheet: The Illuminati Files
  2. Google Spreadsheet: The Illuminati and others
  3. Google Spreadsheet: Some first Organisations
  4. Google Spreadsheet: Events referred to in the Illuminati Files

Our data sets are intriguing and able to attract a wider public interest without further advertisement. They should create steam for the engine if we manage to convinced all the parties involved (Freemason, Berlin State Archive, Gotha Reasearch Centre and the Wikimedia Community) to risk a crowd sourced identification of the roughly 6,000 digitised documents which we have been gathering over the last four years. We have underestimated, however, the problems an empty database (a database without any properties and any primary items) is causing.

If you feel cool with QuickStatements and if you think an empty wikibase installation is just the free space you have been dreaming of, join the team and help us to learn how we can use our data with the brilliant software.

Problem 3: We will need a more massive Wikidata and/or GND input.

We will need our own landscape of information ready to be improved if we want to attract other projects of historical research and regular internet users (with wider a genealogical project for instance). A strategist is here needed, someone with ideas how we could (for instance) acquire all the names of people with birth dates between 1400 and 1800 from Wkidata and/or the GND for our database. (To keep the database clean we might focus on basic data like names, birth dates, places of birth and death, and family connections). Wikibase fans who feel you could organise such an import, feel inspired! We would offer you all the freedom of the experiment you would ask for.

Problem 4: We will need something like forms which users can fill in in order to create standardised CVs with the Wikibase software.

Adrian Heine has taken the first steps into this project. Our aim is a Wikibase environment which normal people can correspond with like they have been corresponding with the Wikipedia software so far. You pick a person of your interest and you get a questionnaire with modules (on places and addresses that the respective person has lived, on employments he or she has been in, on the person’s genealogy, on personal contacts we can prove). Wikibase is presently not exactly ready to be edited by normal people.

Problem 5 (a project for the future): Wikibase needs something like a standard Wikibase-Interpreter

Magnus Manske’s Reasonator has been the cool thing on all my presentations of the Wikibase software in DH-circles. You can pick your language and you get an organised data sheet.

Things get difficult if you want to correct or augment the Reasonator’s information sheet; and things get even more difficult if you want to run the Reasonator on your own platform. The development of an immediate interface that produces smooth pages of structured information will be necessary in order to motivate people to gather information for Wikidata (or any affiliate). The Wikibase-Interpreter would be ready to offer the complete knowledge on any field of interest. It would be ready to list all the places a person is known to have visited, all the contacts he or she is known to have had – whether face to face or through letters. Think of the thousands of contacts of the Leibniz’ correspondence – a problem to be solved with pages that give an overview and “more” on the user’s particular request. The Reasonator is, so far not reading Wikidata directly, nor is it part of the Wikimedia software development. Wikidata will need its own Interpreter in order to become an independent source of information – an independent source that also serves all the Wikipedias around the Globe.

We need to change the way we are organising all this

We have been able to offer a couple of grants in 2017 in order to get the project going. We should be able to use the further funding of DH-projects interested in the software and the collaborative platform to inspire if not to fully finance future tools. DH-projects will, however, only risk a cooperation with the FactGrid project and with Wikimedia as the software and data-partner if we manage to offer an attractive show case of what can be done. The Illuminati files are an immensely cool project to begin with. The global interest in these files is huge, everyone has heard of the Illuminati and here you get their most secret files. Visualisations of networks and of the geographical spread of the secret order will find a good test case here. If we should be able to inspire a crowd sourced annotation of all the known documents – that would stir up a global press attention.

We are, however, far from the show case which we could present anywhere at this moment.

How to use QuickStatements to get data into the FactGrid

The video shows how easy it is to get data from regular spreadsheets via QuickStatements into the database.

You will find a handy cheat sheet of spreadsheet commands, to format your data here: https://docs.google.com/spreadsheets/d/1ov0BL_ob2rFhIl24G4BQhnVAdWeUYOAzlRphEgzOtAE/edit?usp=sharing

Our instance, https://database.factgrid.de/ hosted at the University of Erfurt has its own QuickStatements tool to manage the import of data: Access the tool under https://database.factgrid.de/quickstatements/.

If you do not have an account contact olaf.simons@pierre-marteau.com to get one.