At least a make shift solution: The “Julian calendar stabiliser”

My last blog post triggered a couple of responses on Twitter. It seems I touched a problem that will not be solved that easily.

Save dates as Julian on your Wikibase (manually or, with the /J switch, in your QuickStatements mass input) and your Wikibase will be able to handle these dates correctly in any mixed bag of Julian and Gregorian dates. It is nice that the Query Service is able to produce straight timelines out of any such mixed bag, but immensely problematic that you will be quite unable to get the original Julian dates back in regular Query Service downloads. Blazegraph, the tool that is working behind the Query Service, does its job on normalisations of dates, and these are, of course, performed in the superior Gregorian calendar. Wikibase Query Services are hence on their way to produce loads of unprecedented arithmetical Gregorian dates in environments that have been solely Julian so far. We will first be puzzled by dates that strangely differ from those we fed into these machines — we will have to understand that they have silently added days on them to reach their Gregorian equivalents. Handle these artefacts as correct Gregorian dates, though they are without evidence in the historical records — do not feed them as Julian into any Wikibase because that will immediately expose them to the next round of Julian to Gregorian conversions wherever a Query Service will spot them.

SPARQL queries can actually produce the complexity of the Wikibase they are accessing, but that requires quite some scripting skills. Tagishsimon gave the following script that helped him to get well informed dates from Wikidata in this Twitter response:

Bruno Belhoste applied this script in the following FactGrid query, which will be extremely useful in all future searches on our database. The table gives you the birthdays of members of the French Academy — a typical “mixed bag” of dates that shows all imaginable challenges of different calendars and the various precision statements:

Change the parameters and you will get the dates you are interested in with all the information you will need to process a mixed bag of historical dates from the Wikibase of your choice.

A make shift solution: The “Julian Calendar Stabiliser”

We agreed that we have to stabilise Julian dates on FactGrid under these conditions. All Julian dates will be translated to Gregorian sooner or later on our Query Service. Users must, hence, remain able to get the original Julian information side by side with their (secretly Gregorianised) searches. The simple solution is a string repetition of the Julian statement you want to make. The Query Service does not touch strings, chains of characters and numbers; it will give you the original Julian statement which you can use in other contexts as the very dates you saw in your documents:

Johann Sebastian Bach’s birthday with the “Julian date stabiliser” (see it in the data set)

This is not the ideal solution. One would rather like to have a calendar sensitive Query Service that produces dates as stated on your Wikibase; but it is at least a pragmatic stabilisation to keep Julian dates intact in the waves of transformations and deformations which we are likely to witness in the new world of data processing.

FactGrid Goes NFDI

Friday week before last, we received the news that so many working groups had been eagerly awaiting: the 4Memory consortium (of historical studies) will become part of the Nationale Forschungsdateninfrastruktur (NFDI), the German National Research Data infrastructure.

This is exciting news for FactGrid, just weeks before its fifth birthday. We will be acting as an official repository for historical data in the upcoming NFDI structure. German projects can now make a good case that FactGrid is the optimal platform for their data.

NFDI4Memory task areas

Changing the rules of our present research data management

The German National Research Data Infrastructure aims to bring transparency and sustainability to all research fields, from microbiology to computational linguistics. Whether researchers are still collecting data entirely for themselves in private Word documents and Excel spreadsheets, or whether they are working on digital platforms that are more or less designed like conventional books, designed to be read and looked at – they will face new questions in their research grant applications: Do they produce data? Do they correct publicly available data? If so, the new questions will be: How do they make sure that others can actually work with their data? The idea that new information ends in footnotes of books and articles will not convince the funding institutions any longer. A CSV or JSON data file located on a library server will not do either. Linked Open Data is the only data that is easily reusable – that is what Wikidata has made clear. New platforms are therefore needed – platforms approved by the National Research Data Infrastructure.

The DFG that pushed the process has acted wisely. The different research disciplines had to determine how they would respond to its call for action. They had to create or join umbrella organisations in order to submit proposals for further funding. NFDI4Culture was one of the first groups in the German humanities to receive funding; Text+, for all textual studies, was also among the first arrivals, in 2021. The historical studies collective founded the 4Memory consortium and received the green light in the second round on Friday 4th. Funding will start in March 2023. The Gotha Research Center the 4Memory “participant” on behalf of the FactGrid community in this process.

An international resource as part of a national infrastructure?

It took us a while to feel comfortable with the invitation to participate in this process – back in 2020. At that time we had created a little more than 100,000 items with a handful of participants. Wikimedia Germany was our natural partner. The German National Library was the first major player to collaborate with us in a joint exploration of the Wikibase software. FactGrid from the beginning had invited international collaboration, with projects from France, the United States, Spain, Hungary, and Switzerland. Could we risk a nationalisation of the platform?

The project partners on FactGrid were open to the idea: It would benefit everyone to take the step. The process would open doors to important discussions. We could discuss data standards used worldwide and be able to think of international alliances on this new stage.

Our asset? – Wikibase

Following the NFDI debates,we soon understood why we had been asked to join: We were using Wikibase, the software platform that all members of the nascent consortia were discussing behind the scenes as the very software that could build the bridges between the working groups.

  • Wikibase invites cooperation. Its data modelling is uniquely flexible.
  • Versioning of all editing processes enjoys unprecedented transparency.
  • Wikidata demonstrates that seemingly incompatible fields of knowledge can be managed together in a single graph database.
  • Getting data from a Wikibase platform is as easy as it is to put data into it.
  • Wikibase instances can be federated – we can diversify the scenery without using one single Wikibase instance.

FactGrid was ahead of its time. We were running a functional Wikibase platform while other groups were simply proposing to evaluate the option.

And yet still at the beginning

Over the last two years we have more than quadrupled to 457,000 items. FactGrid is doubling almost every year and there is no reason to believe that this will change in the near future. Projects that are presently preparing data uploads are in the scope of the entire current platform; with our upcoming projects we remain on a global trajectory – we are becoming more international, the platform is learning new languages.

The NFDI process comes just in time because, despite all that growth, we are still right at the beginning, and in urgent need of technological development, which is where we put the focus in our 2020 and 2021 grant proposals. We are not alone in this situation. Wikidata, our elder sister, is still in its initial phase – a peculiar statement, given the fact that Wikidata is celebrating its 10th birthday these days with more than 100 million database objects.

Wikidata is massive. It has rocked the library world as a revolutionary development, but despite that it is still an unknown giant hiding somewhere behind the Wikipedia curtain. Nobody has ever spoken of the data-technical Pentecost miracle which Wikidata actually is. The very name of the project has remained hidden: “Wikidata – you mean Wikipedia, don’t you?”

It is understandable that Wikidata has remained a virtually unknown child. There is neither a search tool leading a wider public to Wikidata information nor is this information readable once you have reached it. The SPARQL query service is a nightmare for normal users. Even if you know how to read computer code– which most of us do not–, how do you find out what information the database can supply? Right, by asking your first specific question with knowledge of the content (the very knowledge that you still do not have). One day an internet-savvy user contacted us with the note that our Query Service had crashed. The Query Service seemed fine; I suggested a video call to get an idea of what the man was seeing on his screen – and it turned out that he was looking at the regular search script. “Send it off, press that blue button!” – He did and received the requested data set. “Ah, I had seen this code stuff but thought it was an error message…”

Wikibase needs two enhancements: An attractive search interface as simple as the Google search box (though with an additional advanced search engine and a SPARQL-search option on top) and browsing software that generates information from the Wikibase or, better still, from several combined Wikibases. The present Wikibase query engine leads you right to the item-pages in the default Wikibase presentation mode, where you can then manually correct or amplify information, but no one seriously enjoys the reading experience. Magnus Manske’s Reasonator, Markus Krötzsch’s SQID, Michael Ringgaard’s KnolBrowser, and Bruno Belhoste’s FactGrid Viewer have shown how Wikibase information can be presented: in pages that present their information concise, well structured, fast to access and easy to exploit. So far, however, all four browsers have remained patchwork solutions. They do not amalgamate platform information in greater depth, and (this is the larger issue) they are as yet not coupled to intelligent search engines. The problem is that we have not yet arrived at independent new resources, at resources whose pages are Google landing points, with pages that amalgamate information from various Wikibases such as Wikidata and FactGrid, and that keep their users on the platform – providing in depth information on request, generating visualisations on the spot, offering downloads of information which users have been accumulating on their tour.

We will get multilingual and attractive Wikibase aggregates. They will integrate information from various resources and they will offer this information in any language requested, identical across all the cultural and political divides. The German NFDI will have to create prototypes of such instruments if they should actually federate Wikibases in a new broader research oriented structure, even if that should start as a national structure.

Opportunities and risks

“The General Intelligence Machine.” Art by H. Lanos for “When the Sleeper Wakes” by H. G. Wells (1899), Wikimedia Commons

The time for a broader research data infrastructure is ripe. Researchers are still handling “their” data on personal hard discs; they copy and paste dates from Wikipedia pages when they could have complete data sets ready to download. Data correction remains fortuitous. Do you write an email to the producers of an online catalogue which you have been accessing with the request to correct a mistake? Do you give the correct date in a footnote of your next article and expect librarians (and Wikipedians) to take note of your work? – We need online resources that allow researchers to correct mistakes right on the screen, in real time; and these resources should be the same ones, which users employ to organise their research. Wikibase is the software that can help to make this possible. How will we get there? Wikibases will have to become the go-to scholarly resources to consult; that is when they will turn into the workbench for the very projects that are using their data.

The landscape of NFDI-consortia comes with its own internal risks. We will need resources to do highly specialised jobs: resources to store and mine texts, resources for the machine readable information which we need in order to make 3D reproductions of objects, and we need resources for historical statements. FactGrid is focusing on this latter need. It cannot become the all-in-one service for historical research. We need the services of other consortia and we should offer our particular services to the other consortia wherever they handle historical statements.

The much more delicate risk of fragmentation looms on the international stage: Will the German expert on French history find herself asked to store her data on a German platform since her funding is German – while her French colleagues with whom she shares the research objects will be delivering their data into a French database? We could, of course, harvest information from 150 national research data infrastructures but that will not provide the same experience for those who generate the information. Working on FactGrid you are about to notice when a colleague in France or China adds to your data. You will contact the colleague with a note of delight about the archival sources that had escaped your notice. Wikibases are joint platforms and should be used as such.

The question of a plurality of national research data infrastructures becomes even more thorny as soon as we look beyond the privileged horizon. We need global platforms to provide equal access to research and to the debates surrounding research. Wikimedia has created Wikidata with the explicit aim of having a software compound on which users from all over the world can work together – accessing and expanding the same pool of global information. We, the international scientific community, the heirs of the international respublica litteraria, shouldn’t fall behind the Wikimedia project.

The fact that FactGrid, an explicitly internationally oriented resource, has entered the NFDI4Memory structure is an interesting development – a chance to get more than one National Research Infrastructure on board.

Links


Header image source: Robert Charles Dudley (British, 1826–1909) Interior of One of the Tanks on Board the Great Eastern: The [Transatlantic Cable] Cable Passing Out 1865/66, Watercolor over graphite with touches of gouache (bodycolor) https://www.metmuseum.org/art/collection/search/383834

9 x FactGrid, Coffee Talk Serie an der Universität Erfurt, 14. April – 16. Juni 2022, Donnerstags 13:30

Die Universität Erfurt lud uns ein, im kommenden Sommersemester eine online Coffee-Talk Serie zum FactGrid als kollaborativer Forschungsplattform zu veranstalten. Neun Themenschwerpunkte haben wir ausgesucht. Die Veranstaltungen sollen kurz und für die Mittagspause zum Hineinschnuppern gemacht sein. Lassen Sie sich inspirieren. Wir bieten eine 15minütige Erkundung mit jeweils offener Fragerunde.

Das Link zur Veranstaltung erhalten Sie für eine Mail an olaf.simons@pierre-marteau.com


14.4.2022: Sich beim Forschen über die Schulter sehen lassen? — Isabella Schwaderers Erkundungen zu den Mitgliedern der Schopenhauergesellschaft

Netzwerkverbindungen in der Schopenhauergesellschaft, 1912, 1913

Kann man es riskieren, auf einer Plattform, auf der alle Daten unmittelbar offen zugänglich sind, die eigene gerade erst angefangene Forschung laufen zu lassen? Isabella Schwaderer tat diesen Schritt mit ihren Recherchen zu den Mitgliedern der Deutschen Schopenhauergesellschaft und wird hier Einblicke in die Nutzerperspektive geben. Worauf lässt man sich ein? Was ist praktisch? Was ist unpraktisch? Was riskiert man? Was gewinnt man?


21.4.2022: Was immer eine Aussage finden kann, kann ein Datenbank-Item werden — wie Wikibase funktioniert

Wikibase steht im Ruf, ganz beliebige Information aufnehmen zu können. Auf einer einzigen Instanz kann man Information ganz verschiedener Fächer zusammenlaufen lassen und sie nahtlos über alle Fachgrenzen hinweg durchdringen.

Das Geheimnis liegt in der Flexibilität Tripel-basierter Daten. Wir können beliebige Objekte aufmachen und Aussagen zu ihnen beliebig an dokumentierte Datenstrukturen anpassen. Die Eingabe ist einfach. Komplizierter und offener ist, wie man die Daten danach in ihrer ganzen Vernetzung intelligent auswertet.

Ein Blick in die Datenmodellierung, die FactGrid Sample Searches und den Query-Service.


28.4.2022: Daten in 400 Sprachen verfügbar machen

Jahrzehntelang kämpfte man in der Bibliothekslandschaft um globale Datenmodelle und verbindliche Datenbank-Feldbelegungen in der Hoffnung, auf sichere Standards. All das hat das Wikidata-Projekt in seiner mutmaßlichen Notwendigkeit relativiert mit dem Angebot einer einzigen Ressource, die jede in ihr gespeicherte Aussage jederzeit in über 400 Sprachen verfügbar macht.

Wie das geht, ist im wörtlichen Sinne trivial: Alle Aussagen werden zerlegt in Datentripel von jeweils zwei Objekten und einer Beziehung zwischen ihnen, deren Teile man nun einzeln in beliebigen Sprachen mit beliebig vielen Labeln belegen kann. Tatsächlich können auf einer solchen Plattform Nutzer, ohne noch über eine gemeinsame Sprache zu verfügen, die Daten aller anderen in der eigenen Sprache lesen – eine gewaltige Chance für Projekte, die in Teams über Sprachgrenzen hinweg zusammenarbeiten sollen.

Ein Blick auf das Wikidata-Projekt, seine Software und Mehrsprachigkeit im FactGrid.


5.5.2022: Selbstorganisation über Projektgrenzen hinweg

Das FactGrid arbeitet ohne zentrale Redaktion, die Daten erst einmal ansehen und auf ihre Qualität hin überprüfen würde. Auch gibt es keine “Relevanzkriterien” – keine Kriterien, die festlegen, was für Daten in die Datenbank dürfen. Wir arbeiten mit einer verwirrenden Offenheit, die dafür ganz eigene Grenzen hat: Forschungsprojekte (auch private) stehen für ihre Arbeit extrem transparent ein. Wir sind hier an einigen interessanten Stellen anders organisiert als Wikidata.

Wie das in der Praxis geht, welche Konflikte man einkalkulieren und welche Konfliktzonen man eher nicht fürchten sollte – Erkundungen der Plattform-Architektur und der speziellen Freiräume, die wir in ihr Projekten gewähren.


12.5.2022, unusual time 18:00 CET: 350,000 objects with cuneiform inscriptions or: Data as a universal language — session in English with Adam Anderson, Berkeley

This is perhaps the most exciting FactGrid project at the moment – designed to create and to interconnect objects for all 350,000 cuneiform artifacts that known today. Where were these objects found? What events, what people, what places are noted on these objects? As a Wikibase installation we would serve as a database a wide collectively could work on and edit simultaneously in all its various present languages. Data stored on the platform would link into other databases and they would be uniquely easy to download for further work in all other software environments.

Our discussion was about the sheer quantities of data such a platform might eventually handle – if we went into the very texture of these objects, locating not only pieces of information but in further steps all the characters on all these objects in 3D data of the artifacts themselves.


19.5.2022: Georeferenzierte Objekte — Session with Bruno Belhoste in English

Paris to download (click on the map) / link for the direct table download TSV formatted

Eine Aufgabe, vor der DH-Projekte immer wieder stehen, ist es, Information auf Landkarten zu visualisieren. Räumliche Beziehungsnetze werden sichtbar, Nähe wird greifbar wie der Horizont, den Verfasser mit Korrespondenzen hatten. Soziale Phänomene, etwa die Zusammensetzung der Bevölkerung in verschiedenen Stadtvierteln, lassen sich erfassen. FactGrid-Information ist jederzeit georeferenzierbar. Wir bieten Georeferenzierungen in einem ersten Projekt – Paris to Download – zur beliebigen Nutzung auf der Plattform oder in anderen Software-Umgebungen an. Einige Blicke auf die Projekte und die Software, die hier nach neuen Modulen ruft.


9.6.2022: Genealogie im FactGrid – mehr als nur Väter und Mütter

What came after Robinson Crusoe’s first edition? EntiTree Visualisation

FactGrid-Information lässt sich komplex in externe Projekte hineinspielen. Zwei FactGrid Browsing-Tools stehen zu Verfügung. Es lassen sich jedoch auch ganz andere Werkzeuge denken.

Als überraschend vielseitig verwendbar erweist sich die von Orlando Groppo und Martin Schibel entwickelte EntiTree-App, die Genealogien auf bequeme Art und Weise mehrsprachig sichtbar macht. Spannende ist dass sich mit der EntiTree App auch noch ganz andere genealogische Beziehungen darstellen lassen.


16.6.2022: Die Zukunft im NFDI4Memory Gefüge – oder: Dateninseln zu neuem Leben erwecken

Das FactGrid ist seit 2021 gesetzt, um im geplanten NFDI4Memory-Konsortium der deutschen Geschichtswissenschaften als Wikibase-Instanz zur Verfügung zu stehen. Forschungsdatenmanagement ist hier das Thema. Was geschieht mit Forschungsdaten, die am Ende irgendwie übrigbleiben – gesammelt, um den Arbeitsprozess zu begleiten, doch danach irgendwie nutzlos, indes voller Korrekturen und Einblicke, die zukünftiger Forschung nutzen sollten? Was geschieht mit Daten, die bislang auf einer eigenen Plattform laufen, nachdem deren Förderung endet? Wie kann man Daten langfristig sichern?

Das FactGrid will hier die Ressource sein, die Information kollektiv nutzbar macht und langfristig in Zirkulation und Korrektur hält. Praktische Tipps, wie das gehen könnte.

A Quarter of a Million Items on FactGrid – just a brief reflection

Germany’s national author Johann Wolfgang von Goethe called it a “masquerade in red and white”, but was himself a member (just as he became a member of the Illuminati a little bit later; it made sense to join such organisations and to know from within what they were all about). Freemasonry was in its most idealistic terms an updated edition of the brotherhood of men united under a simple and strikingly anti aristocratic system: the system of the old craft guilds. With their three degrees of apprentice, fellow and master there was no room for privilege of birth. German masonry evolved from the late 1730’s through the 1750’s principally as a system of four degrees, with Scots Master at the apex and the development did not stop there. The chivalric degrees of the 1760’s and 1770’s gave way to increasingly complex systems, overgrowing this initial construct. These high-degree systems claimed roots in the middle ages if not deeper pasts, synthesising Christianity with alchemy, magic, and theosophy. Masonic entrepreneurs travelled through Europe selling secrets which they would convey in extraordinary lodges. What they offered would have been considered heresies only a generation before, and now became a market of esotericism – a market that turned the masonic world into its first framework and distributor. The Strict Observance or Order of the Temple, the masonic high-grade-system founded by Carl Gotthelf von Hund und Altengrotkau in Germany in 1751 was the biggest player on this stage in central Europe – the system of red and white, the colours of the Knights Templars.

Q250000 is the FactGrid item number of Pierre Faesch, a Frenchman, by profession a gold engraver, who settled in Berlin where he and some of his friends eventually founded their own lodge “Indissolubilis”. He was number 274 in von Lindt’s list of the members of the Strict Observance, published 1846 – number 274 of the 1,266 members he could establish.

Josef Wäges broke the quarter of a millionth item with the input of this list on May 10, 2021 at 7:20 (EST). FactGrid became immediately the most interesting environment for this dataset. 180 of his 1,266 records were old acquaintances: members who already had their Q-numbers on FactGrid. But the new data set which anyone can now create on FactGrid is substantially bigger: it lists 1,595 members with interesting overlaps of projects that have been working on FactGrid over the last three years:

Josef Wäges will publish a more detailed article on the dataset in a lavishly illustrated blog post. The links to the Illuminati are perhaps the most interesting thing to explore in this data set. Von Hund’s claim that the Strict Observance had its roots in the order of the Knights Templar had been both immensely attractive and explosive. The heads of the medieval Order had burned on the stake on May 12, 1310 – but the organisation had gone underground and fused into Scotland’s crypto Catholicism, so the story goes, including the idea that the “Pretender” (to the British throne) was the secret leader of the organisation. The Strict Observance soon expanded from Germany to France, Sweden, Italy, the Baltics and Russia. State leaders became Knights of the Order and met in fancy costumes while members like Goethe or Christoph Bode could easily cast doubts on the historical construct. Von Hund died in 1776 without having given the final proof of the legacy. The organisation itself was by that time in financial troubles over plans to create an insurance system for its members on a foundation of factories, which were to be built under command of the Order on the eve of industrialisation, an organisation that was not really established in the world of modern capitalism.

The internal conflicts culminated in the summer of 1782 when the rank and file of the Observance met at their last convention in the resort of Wilhelmsbad near Frankfurt am Main. The alleged history stood in the centre of the debates and tore the Order apart while a new organisation was secretly emerging behind the scenes: the Order of Illuminati, both as an antithesis and also as a potential heir of the entire infrastructure. They too were by 1782 a masonic high degree system, and they infiltrated lodges far more cunningly from below than from above. With the help of young “Minervals” which they tunnelled from below into the lodges of their interest, and from above with the help of masonic functionaries in the Illuminati leadership. The fascinating thing about the Illuminati was that all the bombastic narratives were handled as little more than a Machiavellian façade by those who acted as “Unknown Superiors” in the hidden centre of this organisation.

Our critical mass: strange organisations of the second half of the 18th century

FactGrid is growing fast. We are doubling our numbers almost every year; that is the more superficial message of the Q250000 jubilee. The more complex message will be: We are (thus far) growing particularly well where we reached our particular critical mass. Entries like Pierre Faesh are the almost ideal subject matter for a Wikibase installation. No portrait has survived, we know little about the biography but we can produce some interesting details with far reaching network information. A genealogy software would not be versatile enough to handle such knowledge. A regular Wiki, with its focus on articles to be written, would on the other hand need to be filled with desolate fragments of repetitive information – we do not know enough to write interesting articles about these people. Using a Wikibase we can easily turn the few points of data we have into an asset. If you want to know more about the “Strict Observance” we can offer the sociological details, networks of the members, family ties, knowledge of the organisation and its surroundings: We can list the various organisational ties of these members and we can – theoretically – give a picture of the landscape of Masonic organisations as they grew and changed from the 17th into the 19th century.

Not quite the software of citizen science: Our Gotha specialisation

At an early point we decided to test the software on the wider audience in a local experiment. Gotha is a small town of some 45,000 inhabitants. We could easily give database courses at the Research Centre. The local project developed with mixed success: Gotha’s Archive of the Lutheran City Church embraced the offer of the free database. Heino Richard of Gotha’s genealogical society entered this project and created its biographical backbone with some 20,000 biographical records linked to the archive’s work and to the city’s history. The integrative appeal remained, however, comparatively weak.

The Wikibase conclusion so far, is interesting in the hands of researchers who are delighted about the flexibility they get with this software. The same tool remains opaque in wider use. We will need interfaces for genealogists and archivists to make broad editing easier, and these interfaces will come.

Novels, religious dissidents, medieval codices and Nazi concentration camps – leaving our comfort zone

We are, nonetheless leaving our comfort zone, the zone of late 18th-century biographies, and this is challenging wherever it leads into fields of information without more comfortable background knowledge:

  • Marie Gunreben of the University of Konstanz has started a project on German novels 1670 to 1750. We have widened this project. We should get the European flow of developments into the picture, the exports and imports, the flow of translations and influences across the European borders. The move is an immense theoretical challenge: We are using a software that creates essential notions of sameness wherever it sets a Q-number. The modern English “novel” should, of course, be the modern French or German “Roman”. But the conceptual equivalents do not really lead us back into the early 18th century. The English “novel” was back then what we will today call a “novella”. Robinson Crusoe, if anything, was a “romance” – a spectacular move in 1719 as the romance had just been pronounced dead, finished by the modern novel(la). How should we handle different conceptual developments in different languages? We are experimenting with set language Items and with Q-items that use the modern conceptual frame as an alleged continuum. It remains to be seen how this will work.
  • Lionel Laborie is about to open the long-expected section on Early Modern religious dissent which our present data have been calling for for the last three years. Freemasons, Rosicrucians and Illuminati, quasi-religious associations built upon a new consensus that their members would leave all their confessional controversies aside and focus on a truth beyond. The result was not exactly deism that shined through all the allusions to God as the master builder and supreme architect. It was rather a competition of increasingly eclectic historical constructs of diverse religious dimensions – of heresies in the old terms of the Catholic or Lutheran orthodoxy and these new orthodoxies emerged within this spectrum with different systems that would not necessarily acknowledge each other. If successful we should be able to eventually give a sketch of the changing map – now with a perspective on the biographies that travelled on this map of ever changing options.
  • Isabella Schwaderer already wrote about her project. She mapped the members of the first two years of the German Schopenhauer Society founded in 1912. The project that began as an experiment led to experiments: Isabel Heide and Martin Gollasch introduced a couple of bigger data sets with the prominent prisoners of Theresienstadt, the map of German concentration camps, and the list of German university academics who signed the declaration of allegiance to the new Regime in 1933. These sets have not yet gained a greater depth of information. They were rather created in order to break the ground for new projects that will discover with a look at early 20th-century networks.

Steps into uncharted territories are a challenge on a Wikibase. You want to augment and to interconnect known objects, you want to work on the basis of our collective present knowledge and suddenly you have to create ever new objects that need ever new objects in order to make sense.

The Middle Ages – the new territory where we will see the biggest growth on our course to Q500000

We will enter new fields and Q500000 is already knocking at our doors. Led by Charles Faulhaber the trilingual PhiloBiblon project has decided to fuse their data into FactGrid – 450.000 items of (late) medieval Iberian books and manuscripts. The project will be a test. We might arrive at the conclusion that the global text production deserves its own Wikibase. It might just as well dissolve the present demarcation lines between archives and libraries on the one hand and historical research on the other. Historical information is in its last consequence not much more than an interpretation of remaining textual and documented evidence. We will bring the evidence and the interpretation onto the same platform.

FactGrid will learn Spanish and Portuguese in the course of this project. The PhiloBiblon group arrives as a team of superbly informed people with different specialisations from data management and librarianship to (literary) history. The technical aim will be to create a user interface on the specific material base that will communicate with the database. FactGrid will act here in the background – nothing to regret, rather the model to go for: The model of a single compound of knowledge that serves various projects as the reservoir of broader collective knowledge.

In the middle of technical developments

Wikibase is not yet a widely used software – it has the potential to become this software. The problem is apparent in any imaginable “normal” use case. You search something – but how do you search anything on the SPARQL Query Service? – on a Query Service that expects you to know what you can search and how you would ask for it – without giving you the slightest hint on either question.

You can use the Wiki surface but here again you will be puzzled. What exactly is the message of these Item pages that collect various statements without order and cohesion? Even if you arrive at a complex item like Q133, Christoph Bode, that item will not tell you half of the story – it does not tell you that this man is the author of hundreds of letters stored in this database, and the recipient of as many – who is mentioned in hundreds of other sources the database has registered.

Markus Manske’s Reasonator gave a glimpse of what one could do with a Wikibase such as Wikidata: One could produce well-structured pages of information automatically in hundreds of languages. The Reasonator did not make it into the software package nor is it easy to use on an external Wikibase.

We will get such browsers – not in the singular but in the plural of general and specific purposes and two of these have entered a test phase last month: Bruno Belhoste’s “FactGrid Viewer” and Michael Ringgaard’s “SLING Browser”. Both seem to do pretty much the same job, but they are doing it differently, opening doors into quite different future developments.

Bruno Belhoste’s FactGrid Viewer (you have been using it over the last minutes wherever you followed the Item-links in this article) is drawing its information straight from the database as you see it. Change data on FactGrid and you will see the new situation with the next browser update. You can switch languages. You get a history of your movements on the site and you get an idea of where you are with a specific item as the object is connected to “what links here?” information.

You can implement Bruno Belhoste’s viewer – pure Javascript – on any website anywhere in the world to see your choice of FactGrid data – the solution for projects who want to use the FactGrid database simply as their database without a further interest in the broader platform.

Michael Ringgaard’s SLING Browser works on the basis of the data dump which FactGrid supplies every evening around 21:15 CET. A new edition of the SLING browser’s presentation of information is created every day. The potential is visible in an intricate detail: The Q-Numbers of SLING browser searches are not necessarily FactGrid Q-Numbers (Christoph Bode our Q133 is on the SLING Browser Q213880). If there is information about the same object available on Wikidata the SLING Browser will give it under the Wikidata Q-Number, and this is only the beginning of the upcoming development: We will eventually see pages that accumulate information from various Wikibases – not in a show of serialised harvests but in a single coordinated representation that accumulates information and that marks the differences only where it arrives at disagreeing statements. This is a tremendous step into the world of “federated Wikibases” that will eventually present the best information of specialised platforms that all speak a common language of triple based statements.

Both browsers are part of the FactGrid-menu-structure but not yet the breakthrough to a simple widespread use of our data. The big issue is at the moment the missing search interface. Google will lead you straight into our items – where you will be lost before you understand how you can navigate on such a platform. The two browsers do not give you a better search interface than the input field on the database’s wiki. If you have just the last name of a person and a rough idea of where they lived that will not help you here or there. You will get to the family name without a hint of how to find those who lived with this Family name. Future Wikibase browsers will have to overcome these dead ends of the individual browsing histories; they will need an advanced search to access data in the first place and internal information that shows why the database has listed the particular object. We will see these interfaces becoming available in a variety of technical options and a broad range of integrations over the next few years.

Integrating FactGrid: NFDI-4Memory participant and GND partner project

We have been surfing a wave of success over the last three years – the wave which Wikibase was creating, the software that is about to be used by national libraries worldwide and in “National Research Data Infrastructures” all over the world.

The reason why National Libraries are experimenting with Wikibase platforms is simple: They have all created authority control data to run the various catalogues that use these data. Humans can understand that Death in Venice was written by Thomas Mann, the 1929 Nobel Prize in Literature laureate. Fresh publications under the same name must have other authors of the same name and this is where databases need a superior form of knowledge. They will handle the 28 authors under that name in a combination of unique identifiers (supplied for instance by the German National Library’s GND’s) and specific biographic background information detailed enough to define who is who in this mess. The system has been working well in its various national boundaries but it was difficult to tell who a specific Thomas Mann was on the BnF’s complementary cataloguing system. It is this riddle that Wikidata has begun to solve. Not only does Wikidata interconnect the up to 300 Wikipedia articles that exist on the various language projects on “the same” entities. The respective Wikidata items will also clarify who these people, organisations and places will be on hundreds of external databases – from the GND to the BnF catalogue.

Historians should states these references on all data they are producing (wherever available) since this is the only way for anyone using their data to automatically check who is who in the different sets they are merging.

The easiest thing FactGrid could do is offer simply all the GND items in the basic pool of objects available on the site to link to. Yet the easiest thing will not be the best thing here. As the National Libraries are about to create and to interconnect their own Wikibases we should enter this compound more as a partner than an interested user. “Our” data should profit from corrections made elsewhere in the wider environment. Corrections made on FactGrid should in return enter the global exchange with information about the research that led to these changes.

We are still living in a world in which DH projects are basically transferring their view of the book world into the new medium of the internet. Books have to be quoted as do web projects – so the common logic, that is creating ever new islands of information on isolated web-platforms.

The future is not the web project quoted in a book or by another web project. The future is in data ready to be downloaded and used in ever new environments. We will need authority to control data, to ensure that those who use our data know what they have downloaded, and we will need collective platforms to offer data in an environment in which the augmentation and further development of information can take place.

https://4memory.de/

It was therefore paramount for us to enter Germany’s present NFDI process. The process is on a trajectory of creating research data repositories in all the fields of the sciences and academic studies – repositories, that will eventually present their data under a broader search engine. We have entered this development as a “participant” of the NFDI’s upcoming 4Memory compound (the compound of the studies that are dealing with historical data).

Our present consideration is how to balance such an integration as a decidedly international site. We will need an international board of FactGrid Stakeholders since this is what we have become over the last three years: an international platform using a multilingual software in order to interconnect research across the borders.

FAQ FactGrid – Pourquoi devrais-je utiliser FactGrid pour mon projet de recherche ?

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).

auf Deutsch
in English
magyar nyelven

Qu’est-ce que FactGrid ?

FactGrid est une installation Wikibase – c’est-à-dire à la fois un wiki ordinaire et une base de données que vous pouvez utiliser pour faire des déclarations sur les objets qui vous intéressent – déclarations que vous pouvez ensuite traiter sur de grands jeux de données dans pratiquement toutes les langues.

La plate-forme est gérée par le Centre de recherche de Gotha et hébergée par l’ThULB Iéna. Elle s’adresse à des projets ayant un intérêt spécifique pour les données historiques.

En collaboration avec Wikimedia Allemagne et le GND de la Bibliothèque nationale allemande, nous essayons d’intégrer cette plateforme dans le prochain consortium d’instances fédérées de Wikibase comme ressource pour les « données de recherche ».

Pourquoi devrais-je utiliser FactGrid pour mes propres recherches ?

Le principal argument en faveur d’un compte FactGrid est la flexibilité imbattable du logiciel Wikibase, que nous avons réussi à installer, dans le cadre d’un projet pilote, en dehors de son site principal Wikidata et avec l’aide de Wikimedia Allemagne :

  • Vous recherchez un logiciel qui parle pratiquement toutes les langues – une plate-forme où vous pouvez entrer des données dans votre propre langue tout en permettant à d’autres de les lire dans leur propre langue ? Wikibase est ce logiciel.
  • Vous recherchez un logiciel qui vous permet de coordonner toute une équipe de manière transparente ? Dans Wikibase, c’est aussi simple que dans le logiciel MediaWiki de Wikipedia.
  • Vous recherchez un logiciel de base de données qui peut faire tout ce que les bases de données d’humanités numériques veulent normalement faire : analyses de réseau, représentations cartographiques, recherches croisées complexes, frises chronologiques (dans différents formats) – un logiciel qui se comporte presque comme un langage humain, tout en fournissant des services complets de base de données ? Wikibase est ce logiciel.
  • Vous avez des données provenant de projets antérieurs que vous voulez développer ? Wikibase dispose d’options de saisie automatique à grande échelle.
  • Vous voulez que vos données puissent être réutilisées ? Wikibase permet le téléchargement et le travail avec vos données, aussi bien hors ligne dans Excel qu’en dligne dans le cadre de nouveaux projets.
  • Vous voulez poser des questions entièrement nouvelles pour votre recherche ? Dans Wikibase, vous pouvez lier n’importe quel objet à n’importe quelle déclaration en fonction de vos intérêts.
  • Vous vous demandez ce qu’il adviendra de vos données et de vos outils de présentation une fois votre projet terminé ? Comptez sur une plate-forme sur laquelle vous ne travaillez pas seul et utilisez une licence de données qui permettra à d’autres personnes de continuer à travailler avec votre travail sans aucun risque !
  • Vous voulez poser des questions entièrement nouvelles dans votre recherche ? Dans Wikibase, vous pouvez relier n’importe quel type d’objet à n’importe quelle déclaration d’intérêt.
  • Vous vous inquiétez de ce qui arrivera à vos données et à vos présentations une fois votre financement terminé ? Comptez sur une plateforme où vous ne travaillez pas seul et utilisez une licence de données qui permet à d’autres de continuer à travailler à la fois avec vos données et vos outils !

Si vous recherchez une perspective à plus long terme, c’est ce que nous essayons d’offrir grâce à notre accord de collaboration en cours avec la Bibliothèque nationale allemande. Nous baserons notre plate-forme sur les données GND afin d’en faire aussi un outil grand public, avec pour objectif de devenir un acteur dans le paysage émergent des “installations fédérées Wikibase”.

Pourquoi ne pas utiliser Wikidata dès maintenant ?

C’est une question légitime à poser. Il y aura des projets (qui utilisent principalement des données) pour lesquels Wikidata sera la meilleure plateforme, comme l’Archivführer zur deutschen Kolonialzeit de la FH Potsdam. Reste que les projets Wikimedia (de même que GND) ne laissent pas de place à la recherche originale. Ils fonctionnent sur des “critères de notoriété” qui ne permettent pas la création à volonté d’objets et de relations entre objets innovants que les chercheurs voudraient tester.

Wikidata et le GND se concentrent sur les informations qui ont déjà été publiées ; ils mobilisent des travailleurs non chercheurs qui alimentent leurs bases de données à partir de recherches déjà publiées. Vous ne serez pas autorisé sur ces plateformes à énoncer des “hypothèses de travail” relevant de “votre recherche”. Vous ne pourrez pas créer des objets de base de données dans le seul but de mener sur eux à un stade ultérieur de votre recherche “rien de plus qu’une analyse statistique”.

Dans FactGrid, nous encourageons au contraire l’utilisation de la plate-forme comme un outil de recherche heuristique.

  • Créez des objets de base de données sur la plate-forme, quelle que soit par ailleurs leur pertinence pour une encyclopédie ou un catalogue de bibliothèque.
  • Risquez comme hypothèses de travail des chronologies provisoires en fonction de vos intérêts personnels.
  • Utilisez FactGrid afin de faire des déclarations non conventionnelles et qui n’ont d’intérêt que pour votre projet de recherche – le logiciel vous donne cette liberté.
  • Créez des objets de base de données spécifiques mentionnant votre projet de recherche dans les jeux de données que vous aurez substantiellement modifiés, ce qui vous permettra d’identifier votre contribution lorsque vous soumettrez votre recherche à votre institution de financement.
  • Risquez de nouvelles hypothèses sur la plateforme et indiquez votre point de vue par un numéro d’objet de base de données (servant en quelque sorte de “micro-publication”), afin d’attester de votre inventivité sur la base de données.

FactGrid est gratuit – comment est-ce possible ?

Le logiciel est disponible gratuitement et en cours de développement dans la communauté plus large des projets Wikimedia et au sein des institutions qui entendent utiliser Wikibase dans les prochaines années.

La plate-forme FactGrid est gérée par le Centre de recherche de Gotha sur un serveur virtuel de l’université d’Erfurt. L’URL allemande coûte 36 euros par an en frais de domaine, financés par le Centre de recherche de Gotha.

Tous les outils Wikidata sont à la disposition de nos utilisateurs. Ils comprennent toutes les applications standard dans les projets d’humanités numériques.

En outre, les logiciels et les outils étant open source, vous pouvez faire appel à votre prestataire de service informatique extérieur pour développer l’application spécifique dont vous auriez besoin.

Proposez à la communauté FactGrid vos propres développements d’outils et de présentation, ce sera le meilleur moyen pour que vos propres visualisations continuent à être développées après la fin du financement de votre projet. Si vous visez plutôt des solutions que vous voulez vendre, vous ne serez pas limité par la licence du logiciel. Vous pourrez commercialiser librement tout ce que vous aurez créé sur la base du logiciel ouvert.

Que dois-je faire des demandes de recherche non orthodoxes ?

Wikibase fait œuvre de pionnier dans la modélisation des données. Pour l’essentiel, vous ne créez que des relations entre des numéros Q (ou entre des numéros Q et des dates, des numéros Q et des coordonnées spatiales, des numéros Q et des fichiers média, des numéros Q et des URL).

Le logiciel ignore le type sémantique des relations que vous avez créées- il s’agit là encore de simples numéros P : Q1 – P1 – Q2 est un “triplet”, qui peut tout aussi bien signifier “Jean-Sébastien Bach (Q1) est le père de (P1) Carl Philipp Emanuel Bach (Q2) ” que “Cette lettre que j’ai trouvée dans les archives avec la cote XYZ (Q1) aurait été envoyée de (P1) Munich (Q2)”

Les numéros Q peuvent être attribués à toute espèce d’identité : personnes, documents, événements, idées… C’est vous qui décidez des types des numéros P dont vous avez besoin pour faire les déclarations qui vous intéressent. Vous n’avez pas à définir les objets dans un système de catégories a priori ; ce sont vos déclarations qui ajoutent de la chair aux objets que vous créez au fur et à mesure. Ne vous inquiétez pas si vous n’avez pas de modèle de données dès le premier jour. Faites des déclarations dès que vous en avez envie et voyez comment elles acquièrent la masse critique. C’est alors seulement que vous pourrez juger de la valeur de l’ensemble.

Toutes les déclarations peuvent elles-mêmes être « qualifiées » – “Jean-Sébastien Bach (Q1) était marié avec (P2) Maria Barbara Bach (Q2) à partir du (P2) 7 october 1707 (date) jusqu’à (P3) environ 5 juillet 1720 (date).” Toutes ces déclarations peuvent à leur tour être dotées de références : “cela ressort du (P4) registre paroissial de… (Q3)”, ”cela est indiqué dans (P5) la biographie bien connue de Bach XYZ (Q4)”.

Le système permet à tout moment de proposer des affirmations concurrentes. Elles sont simplement introduites avec leurs différentes sources et peuvent être comparées les unes avec les autres.

En fin de compte, n’importe quelle déclaration en langue naturelle peut être générée avec des triplets de ce genre, mais, surtout, cela permet d’exprimer cette déclaration dans n’importe quelle langue du monde : pour le système, toutes les déclarations ne sont que des liens entre des numéros Q et des numéros P. C’est vous seul qui attribuez aux numéros Q et P des “libellés” et des “définitions” qui leur donnent sens, et cela dans les langues avec lesquelles vous communiquez (le système gère par ailleurs les informations de date et de quantité dans toutes les normes mondiales avec une conversion automatique dans n’importe quelle direction) ; c’est là le secret des plateformes Wikibase, qui permet aux auteurs de saisir les informations dans leurs langues respectives et aux utilisateurs de lire ces informations dans n’importe quelle langue.

Quels sont les outils fournis par le système ?

Les entrées dans la base de données peuvent être effectuées une à une : ouvrez pour cela l’objet-id en question, allez au bas de la page de saisie et cliquez sur le lien “ajouter une déclaration”. Il vous sera alors demandé de saisir la déclaration que vous souhaitez faire. Vous n’avez pas besoin de connaître le numéro P. Il suffit d’indiquer l’objet dans la langue que vous utilisez et de cliquer sur l’auto-complétion qui vous est proposée. La plate-forme utilisera pour vous le numéro P de cette déclaration. Indiquez alors dans le champ qui s’ouvre l’objet de votre déclaration. Le système, là encore, vous proposera, à mesure que vous tapez le texte de votre déclaration, des suggestions de plus en plus précises.

Les entrées dans la base de données peuvent également être créées et enregistrées automatiquement à partir de tableaux Excel ou CSV. (Ceci est le masque de saisie et voici le guide succinct pour le faire).

Les requêtes dans la base de données doivent être formulées sous forme d’interrogations “SPARQL”, un langage de requêtes qui n’est (malheureusement) pas si facile à utiliser, mais qui, au fond, n’est pas plus complexe que les recherches que vous pourriez vouloir effectuer.

Le plus souvent les utilisateurs de SPARQL ne savent pas écrire leurs requêtes dans le code source. Vous pouvez cependant utiliser des modèles de requêtes où sont indiquées les entrées qu’il faut modifier afin d’exécuter votre recherche spécifique.

En outre, si vous savez exactement le type de requêtes que vos utilisateurs doivent exécuter, vous pouvez créer vos propres masques de saisie, comme ceux que vous utilisez dans les interfaces habituelles des bibliothèques en ligne, qui parleront alors SPARQL avec la base de données.

Le système comprend aussi des outils cartographiques, des frises chronologiques, des réseaux, des arbres généalogiques, des graphiques, etc. Vous n’avez pas besoin de télécharger des applications particulières. Vous demanderez à SPARQL de produire la représentation que vous essayez d’obtenir. Le projet Scholia sur Wikidata présente certaines de ces visualisations.

Que dois-je faire si je veux donner mes représentations de données sur ma propre plate-forme ?

Cela ne devrait pas poser de problème technique. Uwe Jung a montré comment l’interface de la FH Potsdam utilise Wikidata comme dépôt de données sans laisser les utilisateurs voir la base de données à laquelle ils accèdent.

Il n’y a rien de mal à utiliser FactGrid comme dépôt externe et à monter son propre projet de recherche sur le serveur de son université d’origine, en y proposant des accès ciblés à la base de données selon un modèle de recherche de son choix.

FactGrid octroie essentiellement des licences d’utilisation des données à CC0 – cela veut-il dire que je renonce à tous les droits sur mes recherches ?

Opter pour la licence Creative Commons signifie essentiellement que vous conservez tous les droits sur l’ensemble de vos données. Mais surtout, la licence CC0 signifie que vos données deviennent librement utilisables et que vous pouvez ainsi réduire le danger de recherches obsolètes à long terme.

Quelques considérations de base : CC BY 4.0 est à première vue la licence que les scientifiques préféreront. Elle permet l’utilisation gratuite des données à la condition que celles-ci soient correctement citées. En pratique, cela fonctionne pour les textes (comme ce billet de blog) ; dans ce cas, on voit clairement comment on aimerait que le texte soit cité : avec une référence à l’auteur, le titre de la publication, le lieu de publication et la date. Mais supposons que vous souhaitiez que vos données soient citées, disons dans une visualisation ? Une lettre envoyée de Paris à Berlin en juin 1753 se réduisant à une ligne sur une carte, comment cette ligne doit-elle être correctement annotée ? Comment voulez-vous être cité si vous n’avez fait qu’améliorer un ensemble de données existantes ? Les licences “share alike” sont encore plus problématiques : “Ces données sont disponibles gratuitement si les utilisateurs ultérieurs les gardent tout aussi libres”. Cela semble être le plaidoyer ultime pour la gratuité. Mais comment un sous-utilisateur peut-il s’assurer que ses sous-utilisateurs respecteront à leur tour votre contrat de licence (surtout si ce sous-utilisateur offre ses données sous CC0) ? Les sous-utilisateurs seront bien avisés de ne pas utiliser de données provenant de plateformes CC-BY ou CC Share-Alike.

Nos entreprises communes avec Wikidata et la Bibliothèque nationale allemande ne nous ont finalement laissé qu’une seule option : rendre nos données aussi librement disponibles que nos partenaires, autrement dit sous CC0, c’est-à-dire sans garantie que les utilisateurs ultérieurs préciseront toujours exactement qui a collecté les données, ni sur ce que les utilisateurs tiers seront autorisés à faire avec ces données.

 
En pratique, la licence ouverte maximale ne signifie pas que les données de FactGrid sont des données sans auteur, bien au contraire. Outre que nous suggérons aux utilisateurs de toujours citer la recherche qu’ils utilisent, nous faisons l’hypothèse que Wikidata et le GND souhaiteront renvoyer à la recherche sur notre plate-forme. Toutes les modifications apportées aux jeux de données sont liées par le système à nos vrais noms, visibles par tous les utilisateurs. Si un jeu de données a été tout particulièrement travaillé par un projet de recherche, vous pouvez l’indiquer dans une note à part qui sera transférée avec ce jeu de données. Vous pouvez également noter votre travail dans le jeu de données lui-même. Enfin, tout le monde peut interroger la base de données pour savoir quels jeux de données ont été travaillés dans le cadre d’un projet particulier.

En fait, les bases de données comme Wikidata ou le GND de DNB sont intéressées à citer la recherche – cela renforce la solidité de leurs données, et FactGrid est dans la position unique de fournir aux deux institutions une plate-forme sur laquelle les gens peuvent faire ce qu’ils ne pourraient pas faire sur leurs grandes plates-formes.

Que se passe-t-il si je veux continuer à travailler avec mes données sur une autre plateforme ?

Puisque vous avez saisi vos données sans restriction de droits d’auteur, vous pouvez travailler librement avec elles sur tout autre projet qui vous intéresse. En fait, nous aimons être “juste un incubateur” pour les données de recherche.

Que se passe-t-il si les utilisateurs de FactGrid se disputent sur une date “correcte” ?

Le logiciel permet de traiter des données contradictoires – ce qui est particulièrement intéressant dans le domaine de la recherche historique où nous disposons souvent de preuves documentaires contradictoires, sans pouvoir être certain de l’information correcte. Les noms sont traités avec des orthographes différentes ; il arrive que les historiens se contredisent.

Le système permet de reproduire la situation contradictoire ; il permet d’étayer les déclarations avec des dizaines de références et dans différentes orthographes.

Des déclarations divergentes peuvent être comparées les unes avec les autres – par exemple, la déclaration qui fait actuellement autorité et les variantes qui ne circulent qu’en raison des diverses sources contradictoires.

Que deux chercheurs arrivent à des résultats différents, voilà qui devrait en général vous intéresser. Le danger majeur est d’avoir fait une hypothèse erronée et qu’un autre projet sur une autre plateforme donne la solution de l’énigme et travaille sur la bonne date sans que vous le sachiez, ou pire, sans que vous ayez même la possibilité de corriger votre erreur des années après la fin de votre projet.

Pourquoi devrais-je risquer la transparence de mon projet dès le début ?

C’est probablement le problème le plus difficile, celui qui empêche actuellement certains projets d’utiliser la ressource que nous avons ouverte. L’alternative est une ressource accessible seulement avec un mot de passe aux membres de l’équipe jusqu’à la date de publication, c’est-à-dire quasiment jusqu’à la fin du projet. Aucun projet concurrent ne peut alors s’emparer des résultats, du moins en théorie. Personne ne peut voir l’erreur par où vous avez commencé et que vous avez ultérieurement corrigée. Personne ne peut voir non plus le travail fourni par les assistants qui entrent les données dans la base, ni l’implication réelle du chef de projet – tels sont les avantages supposés d’un travail non transparent sur une plateforme qui ne sera mise en ligne qu’à la fin de votre financement.

La recherche transparente offre ses propres garanties : si vous trouvez un document révolutionnaire et établissez une connexion décisive, c’est l’occasion d’attacher la découverte à votre nom et à votre projet. Si quelqu’un, demain, fait la même découverte dans les archives que vous venez de visiter, pas de chance pour lui : vous aurez enregistré votre observation avec un lien dans l’historique des versions que vos rivaux ne pourront pas nier.

En même temps, la plate-forme collective invite à coopérer. Expliquez clairement aux autres équipes sur quoi vous travaillez et permettez-leur de vous contacter sur la plate-forme !

Les inconvénients d’un site web prétendument sécurisé et qui ne sera mis en ligne qu’à la fin du financement du projet sont sérieux : Lorsque le projet est publié, le temps des échanges avec les utilisateurs est déjà révolu. Si la mise sur Internet se fait dans les dernières semaines, celles où le projet est sous pression, vous vous trouverez totalement incapable de réagir par des changements plus conceptuels. Et si vous avez mené des recherches uniquement pour la publication d’un livre, que ferez-vous, vous et votre équipe, des données que vous avez rassemblées dans des fichiers Word et des feuilles de calcul Excel ? Personne ne pourra les verser dans des bases de données, car l’harmonisation à ce stade tardif sera un obstacle insurmontable. Votre seul espoir est que des lecteurs de votre livre parcourront toutes vos notes de bas de page pour en tirer des corrections pour nos catalogues de bibliothèque et pour différents projets Wikipédia. Le risque, finalement, est d’avoir un livre sans impact sur la base de données collective et sur les projets d’humanités numériques et qui sera, à cet égard au moins, obsolète dès sa publication.

L’avenir devrait résider dans une nouvelle attitude à l’égard de la base de données publique. Les chercheurs devraient pouvoir corriger et élargir encore cette base chaque fois qu’ils y accèdent. Pour cela, il leur faut une incitation et une sécurité que seul peut leur donner un environnement de recherche dans lequel le travail soit référençable et citable. Pour cela, Wikibase est mieux équipée que tout autre système.

Comment puis-je faire accepter mon projet sur FactGrid ?

La plate-forme FactGrid ne comporte pas de couche profonde invisible. Tout le monde peut interroger la base de données et les requêtes donneront les mêmes informations, que vous soyez connecté ou non. Votre compte d’utilisateur personnel présente simplement l’avantage de vous permettre de passer à votre langue préférée lorsque vous consultez les données et de voir le lien d’édition sur chaque déclaration.

Si vous souhaitez alimenter la plate-forme avec vos propres données et si vous souhaitez y mener un projet, vous devez disposer d’un compte. Les comptes sont donnés sous des noms réels par les administrateurs. Le logiciel fournit un lien “demande de compte”. Vous pouvez également nous contacter par courrier électronique. Les chefs de projet peuvent recevoir des comptes administratifs leur permettant de donner accès aux membres de leur équipe et aux utilisateurs qui les intéressent.

Une fois connecté, vous pouvez saisir des données en masse ou apporter des corrections spécifiques où bon vous semble. Toute entrée sera connectée à votre compte d’utilisateur. D’autres utilisateurs peuvent annuler vos modifications, mais non sans laisser une trace documentée de cette intrusion dans l’historique des versions – visible par le monde entier.

Si vous souhaitez travailler sur un projet plus complexe, qu’il s’agisse d’une recherche familiale personnelle, d’une visualisation unique dont vous auriez besoin pour une communication, ou encore de l’intégration dans la base de milliers de documents que vous auriez rassemblés dans le cadre d’un projet de recherche de 5 ans, parlez-en à ceux qui sont déjà sur FactGrid et à ceux qui organisent la plate-forme. Nous ne serons pas (nécessairement) désireux de signer un protocole d’entente avec vous, mais il pourrait être très intéressant de faire connaître votre projet sur le blog, ainsi que sur l’ensemble de la plateforme. Là où votre travail devient passionnant, c’est lorsque vous modifiez le travail que d’autres ont déjà fait et que vous encouragez les acteurs d’autres projets à adopter les bons modèles que vous introduisez. Il n’est pas indispensable de discuter des modèles de données avec tous les autres utilisateurs, mais cela peut aider, ne serait-ce que pour diffuser votre travail sur la plateforme. Adoptez des requêtes de recherche composées par d’autres, découvrez des visualisations auxquelles vous n’avez pas pensé, obtenez de l’aide sur la plateforme.

FactGridest conçu pour gérer un environnement de recherche excitant que vous ne trouverez pas ailleurs.

traduit par Bruno Belhoste

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).
Jack Kirby, “The Fourth Dimension is a many splattered thing!” from Alarming Tales, 1 (September 1957).

Celebrating FactGrid’s Q100000: Conrad Alexandre Gérard

Silently and without any fanfare we have passed the 100,000 mark on FactGrid! The item in question is a person: Conrad Alexandre Gérard. In a way he is the perfect candidate to mark this occasion. FactGrid has been diving into networks obscure and less obscure, and Gerard travels on both sides of this distinction: The first French ambassador to the United States and a person interested in Mesmerism, the world of miraculous cures based on “animal magnetism” in the 1780s. Our project started with the German Illuminati and spread into Freemasonry. In doing so, it broadened its scope to include France and England. Gerard is again a perfect representative of this outlook: born in Masevaux, France, in 1729 he pursued a diplomatic career that brought him to Mannheim and Vienna and eventually to the young United States of America. If our hopes are fulfilled, we will follow in his tracks and extend ourselves westwards and across the Atlantic over the next year.

Conrad Alexandre Gérard, Philadelphia 1779
The 100.000th item was created by Bruno Belhoste who began to fuse the Francophone Harmonia Universalis database into FactGrid from where the Mesmerism of 1780s and 1790s Paris will now radiate outwards and into the network of its French and continental adherents. J.J.C. Bode brought these spheres into exemplary contact on his journey to Paris in the summer of 1787 – in the journey that marked the end of the Illuminati since he subsequently declined to resume his work as the last active secret superior when he returned home later that year. Mesmerising is the right word – a word derived from Franz Anton Mesmer, who will stand in the centre of the exceptional scene that is unfolding here.

FactGrid was established in order to give historical data a wider outreach and deeper impact. It is living up to its promise. We welcome joint ventures between platforms. We love to give data an additional outreach. We hope that we can give data sets wider connections and place them within unexpected and exciting contexts of research; and we hope that we can bring research teams together on this mission. Wikibase is designed to encourage this mission.

Looking back: we grew much faster than expected

Bruno Belhoste’s (and David Armando’s) project will deserve a longer blog post in 2020; it will take him another two or three months to feed all his data into the database and to be able to offer more substantial and significant insights; the present input is just preparing the basic structures one can then begin to interconnect. The 100,000th item is more of an opportunity to look back and to speak about the future as far as we can see it from our current vantage point.

Collective editing on FactGrid started on June 11, 2018. We began with data from the two Illuminati research projects that have dominated the scene since the late 1990s. The Illuminati will remain a construction site as we hope to bring online the entire “Swedish Box” (the core collection of Illuminati documents as amassed by Bode from 1782 until 1789). Berlin’s Lodge “To the Three Globes” and the Privy State Archive in Berlin have given encouraging signs that they would support the digitisation and detailed cataloguing. We will need three years of public funding for a project of these dimensions.

The Wikibase installation also invited local low-level projects that were open to testing and developing this resource. Would the “citizens” of Gotha be able to work on one and the same platform used by scholars in their wider international projects? Gotha’s City Church Archive was willing to bring its catalogue online on FactGrid. Heino Richard of the city’s Genealogy Association began an enormous project and has already transferred about two thirds of the “Pfarrerbuch”, volume one of the former duchy of Gotha’s pastors, into structured database information. The former duchy’s 140 pastorates are the project’s backbone; more than 2000 pastors filled the positions since the Reformation. The database has linked them to their parents, their spouses, their respective families, and Heino Richard is now on his way to add all the children — work which will eventually comprise over 15.000 data sets. Here is a perfect opportunity for scholarly professionals since we are dealing here with a tight network of families marrying among each other for over five centuries. We will add information about the professions to allow the sociological evaluation of these family ties.

The Illuminati allowed for the database to grow and merge with Freemasonry — after all this was what they were doing in the 1780s: infiltrating lodges in the German speaking territories. Hermann Schüttler had already identified members of some 130 lodges. Christian Wirkner brought his dissertation on Göttingen’s two late 18th century lodges into the database with some 800 biographies which inspired Martin Gollasch to widen the scope with his own research on the beginning of German’s student fraternities. We are now beginning to understand how the “Landsmannschaften” and a wider spectrum of quasi masonic organisations which recruited students in the central Protestant university cities, laid the ground on which the early 19th-century “Burschenschaften” emerged in the years of the Napoleonic wars. Martin Gollasch’s data sets are enriched by information about all the smaller, more intimate circles of friends which traveled through these wider organisations as he has been mining contemporary “books of friends”, the “Stammbücher” in which students collected entries from their dearest friends.

Bruno Belhoste’s and David Armando’s Harmonia Universalis data will widen the spectrum. One thing has already happened with this new project: Bruno Belhoste has effectively turned the entire site into a trilingual project: All our properties are now available in German, French and English. The interface is already speaking Russian and Chinese (and more than 100 other languages), so there remains some work to be done.

We would love to turn Magnus Manske’s Reasonator into the standard — multi-lingual — interface for simply viewing FactGrid data; that, however, will need a bit of more work from different sides.

One of the projects is hibernating at the moment: Tim Herb made it possible to quote the entire (Protestant) Bible on FactGrid down to the level of the individual verses. The central idea was here to connect all the people and places mentioned in the Bible and the Quran and thus to transform the Biblical historicity and genealogy into structured information. The project is daunting: Wikibase allows contradicting statements. How would a database fare with the competing chronologies of modern and early modern historians? Our predecessors saw the Bible with its succinct 6000 years of Universal history since the creation of the world as the ultimate historical source. When they made the comparison to the Greek and Roman sources, it seemed to them as if the pagans had nourished blatant mythologies. How would the project develop if it was widened into the Quran? The Bible and Quran project is an open challenge at this point. Being able to quote the Bible down to the level of verses has, in the meantime, the charm of offering concise intertextual connections: We might develop a new focus on “Early Modern Networks of Religious Dissent”. The protagonists of this scene had their favorites among the Biblical Books — Daniel, the Prophets, the Apocalypse. Collecting the references we should be able to see how the Bible was used by competing groups on their respective missions, so there is potential here.

The shadow of history: All the GND’s Masonic lodges on a map. The former German speaking regions are coming back to life. We have to get beyond this map – and our map should have historical layers.

Our work on Freemasonry is ultimately as much of a challenge: We could theoretically invite lodges from all over the world to map their historical membership lists and their institutional networks of affiliations and systems back in time. FactGrid should allow visualisations of the spread of Freemasonry on timelines and maps. We are presently at about 850 lodges mostly in the former German speaking territories thanks to a test input of GND data, and these data are broadly unconnected so far — an invitation to dream of the far bigger project that could explore Freemasonry as part of early modern globalisation.

…and ahead: a year of massive challenges

2019 was still a year of cautious consolidation. We have grown faster than expected but we remained a platform of projects that worked silently side by side and in a spirit of open-minded friendship, interlocking knowledge here and there. All data on FactGrid is so far hand picked in tremendously time-consuming work. The GND-input should change the work flow and it should invite projects to start right in the middle of publicly available knowledge. Ten million data sets of people, organisations and places will create a landscape ripe for immediate cultivation.

A lot of questions are still open in this project: Shall we include the whole GND in order to operate as a complete DNB-filial project? Our initial idea was to restrict FactGrid to the early modern period but even such self-imposed and somewhat arbitrary limitations were open to later revision. 1900 had been the line we would draw back in 2018. Today we are confronted with ideas to open FactGrid for research on the entire range of data harvested at the German National Library. How could we deal with personal information of living people without the National Library centre that is responding to requests to modify these data? The opposite threat is just as crucial: And how will we make sure that the input (no matter where it ends) does not turn FactGrid into an agglomeration of data that only a few SPARQL-specialists will be able to mine and which we can hardly keep fresh and alive on our site?

We will have to generate a wider community. We will have to advertise the project in the wider field of Wikimedia projects in order to attract fans of open knowledge from Wikidata and from the different Wikipedia history projects. We will need technical help during the input and we will need a community that adopts this mass of data and that transforms the massive mound of date into a vibrant intellectual playground.

As FactGrid is not exactly a grass roots project we will have to make sure that historical research, archives and libraries will see us as a resource and as a site of collective work. If you belong to the wider world of historians, librarians and archivists and if you are interested in big data you should feel challenged by a project that will turn public data into a treasure one can now, all of a sudden, revise, enrich and explore with unprecedented freedom.

The project will need a more solid technical basis on this course. We need an interface to meet the wider public: an interface which anyone can handle without SPARQL. The interface should be multilingual and it will look more like the Reasonator than our present Wikidata-style pages, which want to be edited rather than looked at.

So, some quite daunting challenges ahead – but we expect it to be inspiring to confront and eventually master them. The software is incredibly cool. It has been opening doors to us during its first one and a half years and we have every reason to think that it will continue to demonstrate this potential for the next years. Wikibase is on its way to become the software of a wide network of Wikibase instances and we should try to become a research platform in this network.