2018-06-11/12: Daten in das FactGrid füllen — praktischer Workshop der Gothaer Forschungsstelle Illuminatenforschung

Liebe Erfurter und Gothaer FactGrid Interessierte,

in den letzten Wochen arbeiteten wir im engeren Kreis an der Einrichtung einer WikiBase-Instanz (der Software hinter dem Wikidata-Projekt) auf dem Server der Uni Erfurt: https://database.factgrid.de/

Es gab einige technische Schwierigkeiten bei der Anpassung der Tools an die neue Server-Umgebung zu bewältigen, doch sind wir seit einigen Tagen soweit, dass die Grundausstattung läuft: QuickStatements steht für den massenweisen Datenimport zur Verfügung. Der Query Service läuft, so dass wir SPARQL-Abfragen von Daten hinkriegen. Das Design (Logo etc.) ist noch offen – was im Moment den Vorteil hat, dass alles wie bei Wikidata aussieht. Matti Blume arbeitet daran, die Datensätze des Illuminatenprojektes für den Projektstart in die Datenbank zu füllen.

Vom Montag den 11. auf Dienstag den 12. Juni 2018 wollen wir gemeinsam mit Matti Blume und Sandra Müllrick am Forschungszentrum Gotha einen Workshop veranstalten, auf dem er praktischen Unterricht in der Befüllung der Datenbank erteilen wird:

  • Wie verändert man einzelne Daten?
  • Wie fügt man über QuickStatements Datenmassen in die Datenbank ein? (Auf dieser Seite ein paar Tutorials: https://factgrid-tools.geschichte.uni-halle.de/blog/archives/811)
  • Wie legt man dazu Properties (Eigenschaften) für Items an?
  • Wie organisieren wir unsere Properties am besten (von dem weit größeren Wikidata-Projekt lernend)?

Projekte, die am Forschungszentrum Gotha und an der Uni-Erfurt im Feld Geschichte und den benachbarten Kulturwissenschaften Datenbankleistung suchen, sind eingeladen, hier praktische Einblicke zu gewinnen. Hilfskräfte, die uns beim Betreuen der Ressource in Zukunft zur Verfügung stehen wollen (und Interesse an den Digital Humanities für die eigene spätere Arbeit haben) betrifft dieselbe Einladung.

Wir werden am Montag einleitend die Software vorstellen und im Austausch mit den Anwesenden ausloten, zu welchen Projekten sie sich eignet. In einem zweiten Arbeitsblock wird es darum gehen, vorbereitete Daten in die Datenbank zu füllen und die Arbeit zu überprüfen. Das wird im Wesentlichen auch das Programm für Dienstag sein, nun mit dem Ziel, Unabhängigkeit bei der Arbeit mit der Datenbank zu gewinnen. Die Uhrzeit für den Beginn der Veranstaltung Montag-Vormittag/Mittag wird im Vorfeld bekannt gegeben.

Die Weitergabe dieser Post an Datenbank-Interessierte ist ausdrücklich erwünscht. Für Voranmeldungen bis zum Mittwoch, den 6. Juni wären wir dankbar, auch für persönliche Notizen zur Hardware: Die Software lässt sich vom Laptop (über Eduroam) bedienen. Wir wollen zusehen, dass wir Computerarbeitsplätze im Forschungszentrum für die Teilnehmer frei halten.

Mit den besten Grüßen,
Olaf Simons

[per e-mail distribuiert]

Zeitplan

Montag, 11. Juni, Pagenhaus des FZG, Seminarraum

(13:15 – 13:40) Olaf Simons: Begrüßung. Kurze Projektgeschichte, eingehenderer Blick auf die Datenblättern aus dem Illuminatenprojekt, die das erste Unterrichtsmaterial geben werden.

(13:45 – 14:15) Sandra Müllrick: Eine kurze Vorstellung von Wikidata und der Wikibase Software. Die Datenbank, die Benutzer editieren können. Versionsgeschichte, Transparenz aller Editiervorgänge über Recent Changes. Triples als Statements. Abfragen über SPARQL, mehrsprachige Datenblätter im Reasonator. Was kann eine Datenbank, was ein reguläres Wiki (wie wir es im Illuminatenprojekt hatten) nicht kann?

(14:25 – 14:45) Praxisteil 1: Wie geht man strategisch vor, wenn man Tabellen aus einem Forschungsprojekt vor sich hat und diese in eine Wikibase Datenbank überführen will? Wie legt man Properties an? Wie behält man Überblick über schon existierende Properties? Wie geht man strategisch vor, wenn man einen solchen Datenberg einarbeiten will?

(15:00 bis zum Erschöpfungsbeginn) Praxisteil 2: Koordinierter Versuch, möglichst viel unserer Daten über QuickStatements in das System zu bringen. Wenn man Tabellenspalten in massenweise Triples überführt – wie legt man die Items an, wie die Properties? (Tutorials unter anderem hier: https://factgrid-tools.geschichte.uni-halle.de/blog/archives/811). Aufgaben für Einzelne oder Zweiergruppen.

Gemeinsames Abendessen

Dienstag, 12. Juni, Pagenhaus des FZG, Seminarraum, respektive Hauptgebäude, Besprechungsraum

(9:00 – 10:30, FZG Hauptgebäude, Besprechungsraum) Planungstreffen: Sandra Müllrick Mitglieder des Forschungszentrums und der Forschungsbibliothek. Projektpläne. Welche Entwicklungen können wir aus eigenen Mitteln finanzieren – wo brauchen wir (etwa bei der Suche von Werkvertragsnehmern) Hilfe von Wikimedia, respektive der Community? Welche Entwicklungen würde Wikimedia gerne anstoßen?

(Parallell 9:00 – 11:00) Praxisteil 3: Fortsetzung von angefangenen Eingaben und individuelle Beratung, insbesondere falls Mitspieler eigene Datensätze in die Datenbank bringen möchten.

(11:15 – 12:00) Praxisteil 4: Wo befinden sich die eingegebenen Daten nun? Was kann man mit ihnen bereits machen? Vielleicht finden wir einige interessante SPARQL-Abfragen.

(12:30 – 13:00) Resümee: Ideen, Desiderate, Pläne.

 

Google Spreadsheets and Illuminati Project Data

Dataset Google Spreadsheet Content What could be visualised? Web Form to build
Documents produced in the Illuminati Order Mostly Schwedenkiste, Illuminati materials from various archives and those documents that were already published in the late 1780s Network information. Caution: We are operating with fragmented and highly selective data. Describe an archival document
Illuminati members and others A list of the c. 1350 members of the order (including names belonging into the wider context), mostly research (stated in col. AJ) by Hermann Schüttler (2016). — The geographical spread of the Order on the central European map: col. AB to be matched with col. AG
— Already existing Wikidata and GND data sets linked in cols. M and N.
— Age structure of the members col. AC
— Percentage of aristocracy cols. J-L
Write a CV
The order had a complex grade system that created careers within the order, people could also get into different official positions Give information about an Illuminati career
Illuminati Events Mostly protocols of gatherings of “Minerval Churches” under Bode’s supervision. (Missing: Hermann Schüttler’s information about gatherings of the “Minerval Church” in Frankfurt/Main) Give information about a session (as type of an event)
Organisations
Publications  from: Forschungsliteratur

Bildnachweis

SPARQL — the Query Language

Wikibase installations are – at this moment – best explored with the SPARQL query language. Specialists are able to write queries in SPARQL but this is not what you would do as a beginner. Most people take a look at an example of a query and then modify the example to suit heir needs.

Here just briefly for the beginning a couple of useful links.


Above: SPARQL in 11 minutes. Note: this video is not specifically on using the Wikibase software.

Above: Navino Evans, co-founder of Histropedia (http://www.histropedia.com/), demonstrating how to construct Wikidata Sparql Queries.

Useful First Aid Links

2018-04-23/25, Antwerp — the first Federated-Wikibase-Workshop

Thanks to the initiative of Andra Waagmeester and Daniel Mietchen (who are part of the community that is behind Wikidata incorporating such staggering inputs as the complete human genome) we ventured a first workshop of projects that have begun to use the Wikibase software outside the original Wikimedia/Wikidata environment. The event at Antwerp’s fifteenth-century hospice, now the Elzenveld hotel and conference centre, was generously funded by the European Research Council. The participants came from fields which only this software would bring together: the natural sciences, the humanities, the social and political sphere and Wikimedia. Wikidata, the Wikimedia project in the centre, is about to become the broadest compound in the world of collaboratively produced knowledge. Starting with interconnecting the 290 Wikipedias all around the globe it became able to switch between all their respective languages. Knowing all the equivalents of articles as well as all the unique items which individual communities created on their Wikipedias Wikidata is closer than any comparable compound to theoretically knowing what a father, a cat, a religion or a novel structurally is. The multilingual competence is matched only by the database’s openness to all sorts of items: Wikidata is dealing with bacteria, the Mona Lisa, Chinese politicians, obscure philosophical concepts, mathematical equations, individual genes, geo-coordinates, and all sorts of properties and qualities these items can gain. The software is open to the flexible creation of ever new properties interconnecting the bricks. You can read Wikidata on the Reasonator and you can explore Wikidata with machines. It is open to logic, accessible in ever new SPARQL queries – the search language worth learning.

Why federating specialised Wikibase platforms will be a win/win situation for Wikidata and the emerging sister projects

Wikidata has overtaken the individual Wikipedias as a project attracting masses of specialised data. Anyone can contribute. The input can be done by machines – so why not the complete human genome? It is at the same moment no longer clear whether Wikidata is the ideal place to host such infusions. Should Wikidata risk the input of all the astronomical objects that have been referenced – of billions of stars creating usually little more than a number and the information of the original observation? This is not only a question of the software’s technical abilities; it is far more a question of the communities needed in order to keep the data easily accessible and up to date.

The alternative is the environment of interacting, more or less specialised platforms that use the same software and that exchange data wherever that is of interest. The win/win situation will have more than one dimension:

  • Wikidata is saved from a flood of information which no Wikidata community can offer to keep attractive.
  • The individual platforms can establish their own workflows and professionalised user rights managements.
  • Wikidata would function more as the bridge between various databases, referring to knowledge elsewhere – a uniquely attractive position.
  • The individual platforms would win visibility and sustainability in the compound – with Wikidata as the central supplier and distributor of information used all around the globe.

A wikibase compound to interconnect the emerging community

The workshop was immensely practical. The central product was not a mission statement of a future collaboration with tight pledges and attempts to institutionalise the emerging field. We were far more eager to discuss software problems, development plans, and to share practical experiences.

The two Wikimedia software developers, Adam Shorland and Raz Shuty, entered a sportive hunt for the stream of “tickets” that emerged in the various debates which Wikimedia’s Sandra Müllrick helped to organise in ever new groups.

Instead of the mission statement we created a further Wikibase instance with the primary aim to map all the projects that have begun to use the software. A retrospective timeline came gratis with the register:

Desirables

It was apparent that we had all been facing technical problems. Some of the projects found their own solutions, others tried to engage Wikimedia as the software’s mother. It was clear that we will all learn from each other and we should aim to make the practical know how we are individually gathering accessible to future projects.

Wikimedia’s Wikibase software is more versatile than any other comparable database software on the market: it is prepared to deal with a constant flux of concepts and specifically designed to adapt to open growth but it is not yet addressing “normal” users, users who are not primarily interested in the technical solutions they are getting here. The SPARQL query is a brilliant key to mine information, yet it is not a language a coincidental Google visitor or a professor of history organising his team will be immediately able and ready to use. Projects will face a need to develop conventional interfaces that will then secretly speak SPARQL in its further dealings with the database.

We will face similar needs to develop web forms which regular users can handle to contribute their knowledge.

Wikimedia – this was the encouraging signal Sandra Müllerick, Adam Shorland, and Raz Shuty were giving as a formidable team of problem solvers – is interested in the increased use of its new product. New projects are encouraged to test the software and if that is of use to them: to become platforms in a far wider compound of projects that can do what Wikidata should not aim to do.

Projects creating the new compound should at the same moment embrace the chance to meet, to learn from each other, and to exchange data. Any data exchange will create new perspectives on their respective fields, it will inspire new tools, it will promote their work in the growing environment of data driven research. The Wikibase software will become common over the next decade.

Our conference was a first meeting and an attempt to look across the borders of our respective fields. We decided to stabilise this exchange; a follow up meeting with probably more participants in New York is in the pipeline.

Details & Links

The Participants

  1. Susanna Ånäs (Finnish Name project; Open Knowledge Finland)
  2. Diego Chialva (Policy Analyst, European Research Council Executive Agency)
  3. Davy Cielen (Data Scientist, Contractor at the European Research Council Executive Agency)
  4. Rajaram Kaliyaperumal (LUMC, Leiden The netherlands)
  5. Lozana Mehandzhiyska London (South Bank University & Rhizome)
  6. Daniel Mietchen (Data Science Institute, University of Virginia)
  7. Lyndsey Moulds (Rhizome)
  8. Alexis-Michel Mugabushaka (Policy Analyst, European Research Council Executive Agency)
  9. Sandra Müllrick (Wikimedia Germany)
  10. Nuno Nunes (Maastricht University, Maastricht)
  11. Adam Shorland (Wikimedia Germany)
  12. Raz Shuty (Wikimedia Germany)
  13. Olaf Simons (Gotha Research Centre of the University of Erfurt)
  14. Greg Stupp (Scripps Research Institute)
  15. Elena Toma (Policy Analyst, European Research Council Executive Agency)
  16. Andra Waagmeester (micelio)

Presentations

Links