Job advertisement in new FactGrid project on Polish-Ukrainian history

Dear Colleagues,

As part of our Research Project „Modelling Premodern Ambivalences“ (VAMOD) at the GWZO in Leipzig we are offering a 10month position until end of December 2025. The project is funded by NFDI4Memory and will offer plenty of possibilities to build up expertise and networks within the field of Digital Humanities.

The goal of the project is to transfer research data from my first book into the Wikibase instance FactGrid in accordance with the FAIR principles. The dataset comprises approximately 800 charters with about 1,600 place names and 5,000 personal names for the present-day Polish-Ukrainian border region (“Crown Ruthenia” or “Red Ruthenia”) between 1340 and 1434. During the transfer, innovative approaches to modelling premodern political configurations will be explored, which will be documented through guidelines and best-practice recommendations, thereby making them sustainably available to the historical research community.

If you can, please share widely!

https://www.leibniz-gwzo.de/sites/default/files/dateien/Stellenausschreibung_Wiss.%20MA_DB2_VAMOD_BS_10.01.2025.pdf

Any questions can be directed to me.
Thank you in advance for sharing.
All the best,

Sven (Jaros)







Dr. Sven Jaros
sven.jaros@geschichte.uni-halle.de
Martin-Luther Universität Halle/Wittenberg
Institut für Geschichte
Professur für Osteuropäische Geschichte
Emil-Abderhalden-Str. 26/27
Raum 2.06.0
06108 Halle (Saale)

Are our Wikibase QueryServices about to mess up two millennia of historical dates?

It was in February 2019 at a conference dinner of medievalists in Jena when I was first confronted with the calendar problem which Wikibase had been posing ever since it had digested its first Julian calendar dates. I had given a Wikibase demonstration earlier that day and now I was sitting next to a medievalist who was ready to destroy me: “Wikibase”, he stated, “is a genuine disaster without anyone understanding it.”

I demanded to hear why that should be the case and the man asked me to show him just one medieval date from Wikidata. I had activated my phone and landed on biography c. 1500.

“See that small print?” he asked, “these dates are all noted as Gregorian before 1584.”

The qualifier was indeed peculiar. Why would they set a Gregorian date before 1582 and then mark it as such? “Well, you know, that the Gregorian calendar was only introduced in 1582, do you?!”

Of course I knew. I am an 18th-century person and Britain had introduced this calendar as late as 1752. The reform had by that time to close a gap of 11 days. But I could also point out that Wikibase allowed the fast correction: “You can easily switch between the calendars, and the machine will actually understand the implications on any timeline” I showed him my screen:

The man was in agony: “Too late. Wikidata is already in big shit”. I realised that I was lacking the full astronomical background and that I did not know the story of these peculiar Wikidata redactions.

Why we needed the Gregorian calendar in the first place

Both, the Julian calendar of 46 BC and the superior Gregorian calendar first introduced in 1582, are approximations. A solar year is one circle around the sun whilst the globe is spinning at about 365.2422 revolutions per year — year after year our planet ends its tour with a different slice pointing towards the sun. We are, in fact slowing down, thanks to the friction which the moon’s gravitation is generating in a constant movement of ebbs and tides, but that is another story. 365.2422 turns per year is our present spin more or less exactly but difficult to generate in a procedural long term pattern of constant adaptations.

The Julian calendar, as it was introduced under Julius Caesar in 46 BC, added one day every four years — in the so called leap years — a rule that boiled down to an additional quarter of a day per year. The approximation of 0.25 days against 0.2422 missed its mark just by 0.0078 days per year, less than a hundredth of a day — negligible one might think — but that one hundredth of a day is a day in a hundred years. In a millennium this discrepancy is piling up to 7.8 days, in two millennia to half a month, moving Christmas further and further away from the longest night until we can finally celebrate Christmas and Easter on the same day; and that was why the Gregorian calendar was finally introduced in 1582 with its far more complex regime of leap years:

  • add one day every four years (as you did under the Julian calendar to create a year of 365.25 days)
  • omit every leap year that is divisible by 100 to get a lower number
  • let this leap year, however, happen if it is divisibly by 400 in order to get a year of 365.2425 days.

The Gregorian calendar reduced the aberration to a surplus of 0.0003 days per year — it will now take 3333 years until we need an additional day to be back in tune with the solar year. The Vatican in Rome adopted the calendar on the 4th of October 1582 — jumping over night into Friday the 15th of that year. Christianity, however, was at that point no longer ready to obey a Papal decree. Eastern Orthodox churches stayed on the Julian calendar right into the the 20th century; Protestant territories and realms would decide one by one. Prussia (with its complex ties into catholic Poland adopted the new calendar in 1612 whilst most of the other Protestant territories stayed Julian for the next 88 years. The United Kingdom took the step in 1752. Lithuania, Russia, and Greece were to switch as late as 1915, 1918 and 1923 respectively.

The following map is from reddit:

When Europe switched from Julian to Gregorian calendar.
byu/coneyislandimgur ineurope

…and it is far from getting the full complexity. The following list gives the growing FactGrid table:

Europe was fragmented. Travelling across Germany in 1699, you could date your letters switching back and forth at every customs house on your tour:

Map of the Holy Roman Empire 1648. Wikimedia Commons

How we solved the problem — and created an even bigger one

Wikibase is a bright software. The tools — the QueryService and QuickStatements — are (or were) not immediately that bright, and that caused the mess the medievalist had noted. QuickStatemens, the tool for mass imports, simply did not offer a Julian calendar switch before February 2023. Instead it would mark all dates as Gregorian without asking — which, looking backwards, was not all that bad…

…why could we all live with the erroneous labelling of Julian dates as Gregorian on Wikidata? Because Wikidata was with this negligence basically doing what we all had been doing up to that point.

Johann Sebastian Bach was born on the 21st of March 1685. Germany’s central database, the GND, is stating this date up until now without the slightest remark on the calendar. The date is Julian because Eisenach’s church register was keeping records in the Julian calendar for another 15 years. The composer himself will not have shifted his birthday to the 31st of March in 1700, the year of the great reset. We all ignore the shift and keep copying dates from documents without any interference. Calendar experts might be interested in the “real” day and they can create Julian/Gregorian calendar matches in those rare cases in which they have to create an exact timeline of events with dates of both calendars.

The Wikidata community had been unaware of the problem. The Gregorian label on all the Julian days was foolish, but the input was actually stabilising our historical tradition as the QueryService will not do anything odd with dates that are entered as Gregorian.

I was far from seeing these advantages after my conversation of 2019 and that was why I warned the PhiloBiblon team in 2022 that QuickStatements would label all their Julian dates as Gregorian against all better intentions once they were imported to FactGrid. Charles Faulhaber immediately asked their programer, Josep Maria Formentí, whether he could not take a look into QuickStatements to solve that little problem. Weeks later Josep introduced the /J-switch that is now available to mark any date as Julian in mass inputs:

+ 1751-06-16T00:00:00Z/11/J

You can now feed thousands of medieval or early-18th-century British dates into your Wikibase and your machine will present all these dates in timelines in perfect synchrony with Gregorian dates. This is extremely nice if you are editing a correspondence whose partners were signing their letters under various calendars. Your machine will give you the exchange of letters in their true course.

So why the alarm?

Wikibase is an intelligent software; it brings objectivity into your statements. Feed a Julian date into your Wikibase and that day will be noted as Julian on the Wikibase itself.

Things get messy wherever we retrieve Julian dates from the QueryService, since this is where the production of funny (and eventually of erroneous) dates will be begin. The QueryService will convert all Julian dates into mathematically correct Gregorian dates.

Martin Luther is known to have died on the 18th of February 1546 — under the Julian calendar, that needs not to be stated, and our Wikibase is giving that date without any calendar stamp on it. But ask the QueryService for Luther’s birthday and it will tell you that the church reformer actually died on the 28th of March, 10 days later — a Gregorian calendar date (without indication) (no big issue you might think, now that you know).

And now think of masses of data which we will be moving between Wikibases in the brave new world of “federates Wikibases”. If there are “Julian” dates among them, then these will get secret additional days wherever they are extracted with the help of a regular SPARQL-Query on the QueryService.

This is what will happen to Luther’s date of death as it is now no longer a subject of safe copying. We will see it in an increasing number of variants — namely as:

  • 18 February 1546 (Greg.) — mistaken QuickStatements input artefact
  • 18 February 1546 (Jul.) — the historically correct date
  • 28 February 1546 (Greg.) — unorthodox but correct Wikibase QueryService output
  • 28 February 1546 (Jul.) — Wikibase output mistakenly saved as Julian
  • 10 March 1546 — the previous converted to Gregorian
  • 20 March 1546 — the previous after the next im- and export

and so on and so on.

Can we stop the wave of uncontrolled additions of days on Julian calendar dates?

I am not quite sure how. We need a QueryService that will never ever offer a historical date without the corresponding calendar statement (now that we have a machine that does both calendars).

But not only the QueryService is posing a problem here. Our Wikibases should have a third option, because our documentary evidence is usually lacking calendar information. Eisenach’s church register of 1685 is using the Julian Calendar (without further notice), that is something we can determine — but we cannot say what calendar an author of a typical letter was using in 1685 if that date comes without a localisation. Our documents do not tend to have calendar statements on them.

What we need here is a third — a “calendar format unknown” — option. It’s complicated, I am afraid.

Links and more

  • Header image from Ολυμπία δώματα, or, An almanack for the year of our Lord God 1752 (London: Printed by T. Parker, for the Company of Stationers, 1752), from the digitisation at Archive.org
  • English Wikipedia List of adoption dates of the Gregorian calendar by country https://en.wikipedia.org/
  • See also: Maniphest T207705, Implement the Extended Date/Time Format Specification, https://phabricator.wikimedia.org/T207705
  • Lydia Pintscher, calendar model screwup, 30 Jun 2015. [https://lists.wikimedia.org/hyperkitty/list/wikidata@lists.wikimedia.org/thread/Y7OEHUYV66DHRVZ6JCSODWAYZ25SLUHM/ https://lists.wikimedia.org/hyperkitty]
  • Julian and Gregorian dates from Wikidata, question asked on https://opendata.stackexchange.com/, Apr 18, 2018 at 0:33 [https://opendata.stackexchange.com/questions/12723/julian-and-gregorian-dates-from-wikidata https://opendata.stackexchange.com/]

Erste Hilfe beim Zuordnen mittelalterlicher Ortsnamen (5770 Vorschläge)

Anfang des Jahres fragten wir (ich gab die Frage für Kathleen Schnabel und das Team Robert Gramsch-Stehfests ins Netz) die Welt der “Twitter Mediävisten” nach einem klugen Tipp, wie wir gut 3000 mittelalterliche Ortsnamen identifiziert bekämen. Es handelte sich um Ortsnennungen, die Studenten, die sich zwischen 1392 und 1450 an der Uni Erfurt einschrieben, zu ihren Namen in die Matrikellisten gaben, niedergeschrieben wohl immer nach Gehör.

Der Tweet war erstaunlich erfolgreich: 13.900 mal gesehen, 115 mal geliked, 100 mal weiterversandt. Hilfreiche Antworten kamen aus allen Richtungen.

Offenbar waren wir nicht die ersten, die an diesen besonderen Abgrund gerieten.


Aberwysczel
Abswinden
Adelenessen
Adelfessessen
Adenstede
Adernheym
Adirstete
Aemstelredam
Agghelbeke
Ahorn
Ailsfeldia
Akusgrann

Natürlich hätten wir einfach bei den Immatrikulationen, die wir verzeichneten, in einem eigenen Feld notieren können, was die Studenten als ihre Herkunftsorte angaben, respektive die Schreiber daraus machten. Das taten wir auch am Ende. Wer aber von den Studenten eines Jahres aus demselben Ort stammte? Wo das Einzugsgebiet der Uni lag? Wie es sich mit dem Aufstieg der Uni veränderte? – das alles ließ sich ohne Identifikationen der Orte nicht klarer ermessen.

Banal war es, die Ortsangaben in einem Google Spreadsheet allen Orten der Datenbank gegenüberzustellen und mit VLOOKUP (SVERWEIS) eine unimittelbare Zuordnung durchzuführen. Die Treffermenge fiel aber unbefriedigend schmal aus und war durchzogen von sich auftuenden unterschiedlichen Problemen. Gotha konnte in den Matrikeln als Gota oder Gotta auftauchen, nur ein einziger Buchstabe verhinderte in diesen Fällen das “matching”. Bei Namen wie Akusgrann lagen die Dinge dagegen komplexer. Hier sollte man wissen, dass der Ort lateinische auch als Aquisgranum und Aquae Grani bekannt war und so zu Eindeutschungen verleitete.

Michael Markert von der Thüringer Landesbibliothek Jena schlug mit einem YouTube Tutorial den Abgleich vor, der das Feld der Treffer in dieser misslichen Lage unmittelbar handhabbarer machte – eine GND/Lobid Anfrage, die einen mathematischen Buchstabenaustauschverfahren Varianten ins Kalkül brachte, mit denen sich alle geringfügigen Schreibunterschiede erst einmal auflösten:

Skurrile Treffer machten die sehr speziellen Schwächen dieses Angebots deutlich: Argentinische Studenten wollten sich da 90 Jahre vor der europäischen Entdeckung Südamerikas in Erfurt eingeschrieben haben, Studenten „de Argentina“, aus Straßburg.

Ein eigenes Problem blieben zudem die Orte mit aktuellen Namensgleichheiten. Dem Abgleich fehlten Wahrscheinlichkeitsparameter, Formen eines eigenen Kontextes: Von zwei Rothenburgs sollte das bei Fulda eher im Einzugsbereich der Erfurter Universität liegen als das an der Tauber oder das an der Wümme (entscheiden ließ sich das letztendlich jedoch nicht). In anderen Fällen war eher über Infrastrukturen nachzudenken: Wenn es zu einer Nennung ein Dorf und einen Ort mit mittelalterlicher Lateinschule gab, war vermutlich eher der Ort mit der Lateinschule der Entsender.

Am Ende blieb nichts übrig, als in einer Gruppensitzung alle Vorschläge zu überprüfen und bei vielen der Angaben historisches Wissen spielen zu lassen – bei 3000 Ortsnennungen ein gerade noch gangbarer Weg.

Das Endergebnis blieb eine Annäherung und erweist sich im Moment als vorurteilsbehaftet: Österreich, die Schweiz, die Niederlande sowie die ehemals deutschsprachigen Ostgebiete dürften in der folgenden Karte unterrepräsentiert sein (man kann in das Iframe hineinzoomen, die Karte wird bei jeder Browserauffrischung frisch aus den Datenbankeinträgen generiert). Was hier sichtbar wird, ist, dass wir Orte nach heutiger Nationalität gebündelt in die verschiedenen Schritte des Abgleichs brachten, um dabei annäherungsweise räumliche Nähe ins Spiel zu bringen.


Zoomfähige Kartendarstellung: Die Herkunftsorte der Erfurter Studentenjahrgänge 1392 bis 1450

Erst Blicke in die Biographien werden die Entscheidungen substantiieren können. Die Datenbank erlaubt es indes, Baustellen aufzumachen. Dies ist die Liste aller Orte, die wir im Moment für eine “manuelle” Überprüfung zurücklegten:

Doch ein nützliches Produkt am Ende

Dem Bedarf, der sich hier auftat, Rechnung tragend, spiegelten wir die durchgeführten Identifikationen am Ende auf die heutigen Ortsnamen zurück, so dass sie sich nun zwei neue Handhabungen ergeben:

Gibt man im Suchschlitz des MediaWikis einen Namen ein, den man nicht sofort einem heutigen zuweisen kann, so erhält man mögliche Treffer unmittelbar über die Alias-Funktion der Wikibase-Instanz angezeigt.

Ortszuweisung nach den Aliasangaben über den einfachen Suchschlitz

Spannender aber sollte die Liste unserer Zuweisungen sein. Sie lässt sich nun unmittelbar aus dem folgenden Fensterausschnitt als CSV, TSV oder JSON-Datei herunterladen (die Download-Links erscheinen am rechten Fensterrand im Mouseover; Quellennachweise finden sich in den einzelnen Datensätzen und können mit einer komplexeren Suche auch hinzugeladen werden):

Als TSV Datei heruntergeladen lässt sich die nun spaltenweise erscheinende Liste jeder eigenen in einem Excel- oder Google-Datenblatt gegenüberstellen und mit VLOOKUP/SVERWEIS auf einfache Art innerhalb des Datenblatts abgleichen.

Weihnachten rückt näher. Wunderbar wäre ein Tool, mit dem man Abfragen kontextualisieren könnte. Wir suchten Orte aus dem Spätmittelalter mit einer Fokussierung auf Erfurt und einer Privilegierung von Orten, die im Mittelalter über Schulen verfügten. Nicht einfach. Die Macher des GOV, des Genealogischen Ortsverzeichnisses sollten hier viel weiter sein – vielleicht dass wir einmal zusammen einen viel intelligenteren Service zu Ortsnamen auf die Beine stellen.



Publiziert im Rahmen des der NFDI4Memory Task Area “Data Connectivity”, Historisches Datenzentrum Halle, Projektnummer 501609550.

PhiloBiblon receives a new grant from the National Endowment for the Humanities

We are delighted to announce that PhiloBiblon, a database of the primary sources for the study of medieval Iberia,  has received a two-year implementation grant from the Humanities Collections and Reference Resources program of the National Endowment for the Humanities to complete the mapping of PhiloBiblon from its almost forty-year-old relational database technology to the Wikibase technology that underlies Wikipedia, Wikidata, and FactGrid. The project will start on the first of July and, Dios mediante, will finish successfully by the end of June 2025.

The fundamental problem is to map the 422,000+ records of PhiloBiblon’s bibliographies with their complexly interrelated relational tables to the triplestore structure of Wikibase. A triplestore relates two Items by means of a Property. Thus a Work is linked to an Author by the Property “written by.”

We received an NEH Foundations grant for this project in 2021, as described in detail in PhiloBiblon: From Siloed Databases to Linked Open Data via Wikibase: Proof of Concept. Over the course of the last two years, the pilot project team, consisting of Charles Faulhaber (PI), Patricia García Sánchez Migallón and Almudena Izquierda (doctores por la UCM), Berkeley undergraduate Spanish and data science majors (Julieta Soto, Serena Bai, Tina Lin, Cassandra Calciano, Martín García Ángel), Max Ziff (data engineer), and Josep Formentí (user interface programmer), with the guidance of Olaf Simons, has analyzed the data structures of PhiloBiblon’s ten relational tables (using BETA for the test cases) and worked out the procedures needed to convert them into triplestore structures.

Almudena and Patricia manually mapped more than 125 BETA records to FactGrid: PhiloBiblon as models for the automated processing of the rest. See for example the records for Alfonso X, BNE MSS/10069 (Cantigas de Santa Maria), and the 1497 edition of the translation of Boccacio’s Fiammeta. These models have been key for establishing the semantic relations between PhiloBiblon’s data fields and the Properties and Items in FactGrid. In many cases appropriate properties did not exist and it was necessary to create them. For example, something as simple as the Watermark property was needed in order to identify the various watermark types set forth in PhiloBiblon’s controlled vocabulary.

Julieta Soto and Martín García Ángel attacked the problem of creating almost 900 FactGrid records for the controlled vocabulary terms in BETA. This meant in the first place a search in FactGrid to make sure that an equivalent term did not already exist, in order to avoid creating duplicate records. Then they had to situate the term in the FactGrid ontology by specifying it as a “basic object” (e.g., fruit) or identifying it as a subclass of an appropriate basic object, for example facsímil impreso as a subclass of facsímil. At the same time they had to link the record to the code in PhiloBiblon, BIBLIOGRAPHY*RELATED_BIBCLASS*FAP, identifying a record in the Bibliography table as a print facsimile, thereby making it possible to search for such items.

The default viewer used in FactGrid, the same as that used in Wikidata, is not user friendly. Therefore Josep has created a prototype user interface, using data from the BETA Institutions table. We encourage you to play with it and tell us what you like or—more usefully—don’t like.

This change to Wikibase technology is designed to allow PhiloBiblon not only to take advantage of the linked open data of the semantic web but also, and most importantly, to decrease sustainability costs. Because Wikibase is open-source software maintained by WikiMedia Deutschland, the software development arm of the Wikimedia Foundation, software maintenance costs for PhiloBiblon will be minimal in the future. This means that it will no longer be necessary to seek major grant support every five to seven years merely to keep up with technology change.

While this work has been going on, we have not neglected the vital process of cleaning up PhiloBiblon data in order to facilitate the automated mapping nor the equally vital process of adding new information to PhiloBiblon. For example, Pedro Pinto, a member of the BITAGAP team, has recently discovered a “folha desmembrada” (BITAGAP manid 7862) from the Livro 4 of the chancery records of king Fernando I (1345-1383) (BITAGAP manid 3255), separated from the manuscript in the Arquivo Nacional da Torre do Tombo. The newly discovered dismembered leaf contains five previously unkown royal documents. It was being used as the cover of the “Livro de Acordãos, 1620-24,” in the archive of the Santa Casa de Misericórdia in Coruche, a small city in the Santarem district on the Tagus river northeast of Lisbon.

The recycling of parchment leaves from discarded medieval manuscripts, presumably for more socially beneficial purposes, such as the protection of administrative records, was common in both Spain and Portugal in the sixteenth and seventeenh centuries. Such leaves have been the source of many previously unrecorded medieval texts. Perhaps the most spectacular exemple was Harvey Sharrer’s discovery in 1990 of the eponymous Pergaminho Sharrer (BITAGAP manid 1817), with seven unknown poems of king Dinis of Portugal (1279-1325). This had been used as the binding of a collection of notarial documents (Lisboa: Arquivo Nacional da Torre do Tombo: Lisboa, Cartório Notarial de. N. 7-A, Caixa 1, Maça 1, livro 3).

PhiloBiblon: From Siloed Databases to Linked Open Data via Wikibase: Proof of Concept

We are very pleased to announce a pilot project funded by the U.S. federal government’s National Endowment of the Humanities (NEH): “PhiloBiblon: From Siloed Databases to Linked Open Data via Wikibase: Proof of Concept.” It will begin June 1, 2021. and end May 30, 2022 and will be hosted by FactGrid.

The project is focused on PhiloBiblon. a forty-year-old database for the study of the medieval history and literatures of the Romance cultures of the Iberian Peninsula: Portuguese and Galician-Portuguese, Castilian, and Catalan. It contains four subsidiary databases:

  • BETA: Bibliografía Española de Textos Antiguos: Medieval texts in Spanish.
  • BIPA: Bibliografía de la Poesía Áurea: Golden Age (16th-17th c.) Poetry in Spanish.
  • BITAGAP: Bibliografia de Textos Antigos Galegos e Portugueses: Medieval texts in Galician, Galician-Portuguese, and Portuguese,
  • BITECA: Bibliografia de Textos Antics Catalans, Valencians i Balears: Medieval texts in Catalan.

The project is designed to solve one of the most vexing problems facing long-standing digital projects: maintainance of the software platform. PhiloBiblon started out in 1975 as an ancillary database of the Dictionary of the Old Spanish Language project, carried out at University of Wisconsin, Madison, by Lloyd Kasten and his student, John Nitti. In Madison the database management system (DBMS) used was FAMULUS, created, ironically, in Berkeley in 1964, for the Pacific Southwest Forest and Range Experiment Station. Since then, technological transformation has been a constant: from the CD-ROM discs of ADMYTE (Archivo Digital de Manuscritos y Texts Españoles ) to a first web version in 1997 to the current 2014 version. Each transformation has usually required multiple grant applications to agencies and foundations, especially to NEH.

PhiloBiblon currently runs under Windows on Revelation Technology’s MultiValue database, OpenInsight. PhiloBiblon’s 1987 implementation was designed on Revelation G., the ancestor of OpenInsight, by John May, a graduate student in History of Science who eventually left the academy for a business career. John has maintained and enhanced the PhiloBiblon DBMS for more than 35 years, building it out to encompass ten relational tables (texts, witnesses, primary source manuscripts and imprints, copies of imprints, persons, institutions, toponyms, and secondary references). Among them these ten tables contain 1246 data elements (fields), 98 controlled vocabulary lists with more than 3000 properties, 110 search indexes, and 30 data entry screens.

PhiloBiblon and its relational DBMS existed before Tim Berners-Lee invented the Worldwide Web at CERN in Geneva in 1989. It was evident from the advent of the first really useful commercial browser, Netscape, that the WWW offered a vastly superior vehicle for making information available to students and scholars, much better than print and CD-ROM. In PhiloBiblon’s first web version (1997), data were exported from the OpenInsight DBMS and uploaded to a web server at Berkeley in HTML format. Users could search only for authors, titles, or keywords; and the result of the search was just a list of manuscripts or editions, each of which had to be opened, one-by-one, to find the item of interest.

In the current version OpenInsight data is still exported—usually every two or three months—and uploaded to a web server at Berkeley (and to a mirror site at the Universitat Pompeu Fabra in Barcelona) in ten separate files, one for each table, in the more advanced XML format. On the server the eXtensible Text Framework (XTF) program, created by the California Digital Library of the University of California, parses each file into individual records, indexes them, and serves them to users as a result of a search request from the PhiloBiblon web site. The search mechanisms are much more powerful, offering not only keyword searching but also a set of search boxes tailored to each entity.

Thus for texts (works): keyword (simple search), author, title, incipit, explicit, associate person, date and place of composition, subject.

For manuscripts and editions: City and library holding the item, shelfmark, date and place of production, printer and publisher, scribe and patron, previous owner or other associated person:

The time lag between data input and its appearance on the web is annoying. Moreover, the process required to export data and upload it to the web to the web is is neither elegant nor efficient. Aside from this and from the problem of maintaining and enhancing both the Windows DBMS and the web software, the most urgent issue facing PhiloBiblon is its status as an information silo, with no organic relationship with other information sources. The web 3.0, the semantic web, is designed to make use of Linked Open Data and the Resource Description Framework (LD/RDF) in order to make possible automatic links to other information resources, like the Virtual International Authority File (VIAF).

Since 2014 the PhiloBiblon research teams have prepared a series of unsuccessful grant proposals, separately as well as in collaboration with other projects, to NEH and to Spanish, Catalan, and European agencies and foundations with the goal of funding the transformation of PhiloBiblon into an LD/RDF resource.

Last year, instead of proposing the creation ex professo of a new web-based DBMS for PhiloBiblon, on the advice of our neighbors at Stanford University, the University of California, Davis, and the international library consortium OCLC, we decided to explore a radically different solution: to incorporate PhiloBiblon into the wiki world. PhiloBiblon already cites Wikipedia constantly, more than 1400 times in nine different languages. Of more interest than Wikipedia as a model, however, is the more structured but still open environment of Wikidata. Wikidata, however, is too open. PhiloBiblon requires more control over the individuals who can contribute to it.

This led us to FactGrid, which is ideal for our purposes. Open only to members, it offers a perfect sandbox for the PhiloBiblon staff to explore the relationship between PhiloBiblon’s highly structured data model and the elegant and infinitely extensible model of triplestores based on Q# entities and P# properties. We are enormously grateful to Olaf Simons not only for his generous offer to make it available for this purpose but also for agreeing to serve on the Advisory Board of the current NEH project.

When this project ends May 31, 2022, we hope to have shown that the Wikidata model is viable for PhiloBiblon over the long term and that we can make use of its standard input processes, modified as necessary, to map PhiloBiblon’s 421,000 records into the corresponding FactGrid entities and properties, creating new ones as necessary. This work will be carried out primarily by data analyst Adam Anderson, whose academic specialization is Assyriology and cuneiform studies. He will also study Wikibase’s LD access points to and from libraries and archives and test the Wikibase data export module for JSON-LD, RDF, and XML on PhiloBiblon data.

TABLA BETA BITAGAP BITECA BIPA
ANALYTIC (witnesses) 14692 52084 12239 89926
REFERENCES 7270 21558 5976 472
PERSONS 7423 32309 3473 3418
GEOGRAPHY 1814 4759 840 211
INSTITUTIONS 794 3297 585 4
LIBRARIES 915 455 420 119
MANUSCRIPTS & IMPRINTS 5168 5886 1971 1572
COPIES OF PRINTED BOOKS 4157 1146 1473 137
SUBJECT HEADINGS 339 34 149 126
WORKS (texts) 6034 31962 6173 89913
TOTAL 48606 153490 33299 185898 421293

In addition, software engineer Josep María Formentí (Barcelona), after evaluating the Wikibase data entry module and report format, will create prototypes of more user-friendly query and data entry screens and report formats.

All of this work will be carried out in collaboration with the twenty members of the PhiloBiblon volunteer academic staff and, we hope, numerous volunteers from the Hispano-medievalist community.


Image: Rueland Frueauf the Elder (1440–1507) The Education of the Infant Christ (1506) Wikimedia Commons.

Ein Best-Practice-Szenario für die Erschließung historischer Wissens- und Gebrauchsliteratur als Open Data

In Wikiversity erstveröffentlichter Wikimedia-Wettbewerbsbeitrag

Projektbeschreibung

Das Wissen über die Welt und den menschlichen Umgang mit dieser, über praktische Fähigkeiten und theoretische Erkenntnisse, wurden über Jahrhunderte in Handschriften gesammelt und verfügbar gemacht. Die ältesten deutschen Wissens- und Gebrauchstexte stammen noch aus Althochdeutscher Zeit (8. – 11. Jahrhundert) und besonders in Frühneuhochdeutscher Zeit (ca. 1350¬1650), wächst die Anzahl der Texte und Themengebiete rapide. Immer neue Wissensbereiche wurden in deutscher Sprache erschlossen und ein Großteil der spätmittelalterlichen deutschen Handschriften enthält Wissens- und Gebrauchstexte. Dennoch stehen diese Texte nicht im Zentrum germanistischer Forschung und sind – auch aufgrund ihrer Diversität und Komplexität – wesentlich schlechter erschlossen als literarische Texte. In meiner Forschung versuche ich diese Wissenslücke zu schließen und bislang vernachlässigte Textsorten, wie etwa Losbücher, Kalender, Geomantien, Textamulette, Tintenrezepte oder Anleitungen zur Dämonenbeschwörung (auch hier, hier und hier) so zu erschließen, dass sie von Fachkollegen, aber auch einem größeren Publikum, aufgefunden, gelesen und verstanden werden können. Dies soll auch dazu dienen die Überlieferung dieser Texte und damit die geographische wir soziale Verbreitung historischer Wissensbestände nachvollziehbar zu machen. Ein großes Problem ist dabei die mangelnde Referenzierbarkeit der Texte, die aufgrund ihrer Form (bspw. Rezepte) oder ihre Unbekanntheit keine etablierten Titel haben, in den einschlägigen Fachlexika (Verfasserlexikon) nicht erfasst sind und auch innerhalb der Forschungsliteratur unterschiedlich bezeichnet werden. Im digitalen Umgang mit diesen Texten verstärkt sich dieses Problem, da handschriftlich überlieferte deutsche Texte überhaupt nur in Ausnahmefällen über Normdaten oder andere Linked-Open-Data-Formate erschlossen sind.

Im Rahmen des Fellow-Programm Freies Wissen soll ein Best-Practice-Szenario für die Erschließung historischer Wissens- und Gebrauchsliteratur als Open Data entwickelt werden. Dieses schließt an bisherige Arbeiten zu den Losbüchern und Chiromantien sowie der laufenden Katalogisierung der illustrierten mantischen Prognostiken für den Katalog der deutschsprachigen illustrierten Handschriften an und soll die deutschsprachigen mantischen Texte des Spätmittelalters einem breiteren Publikum erschließen und Forschungsdaten nachnutzbar machen. Dazu ist eine Kombination aus Open-Access-Forschungsbeiträgen, der Überarbeitung von Wikipedia-Artikeln, der Veröffentlichung von Handschriftenabbildungen (im Archive und in Wikimedia Commons) sowie der Generierung von offen zugänglichen Erschließungsdaten vorgesehen. Dabei sollen Erschließungsdaten zu deutschsprachigen Wissens- und Gebrauchstexten des 14. bis 16. Jahrhunderts erstmals in der FactGrid-Datenbank und damit auf einer Wikibase-Instanz erfasst werden. Über die FactGrid-Datenbank sollen nicht nur einzelne Texte dauerhaft und eindeutig identifizierbar gemacht, sondern auch verschiedene digitale (Handschriftendatenbanken, Bibliothekskataloge, Handschriftendigitalisate) und analoge (Forschungsliteratur) Angebote vernetzt werden, sodass Informationen zu einzelnen Texten zentral auffindbar sind. Gleichzeitig ermöglicht die Erfassung von Textereignissen einzelner Handschriften und Werken in FactGrid auch eine maschinelle Auswertung und Visualisierung dieser Daten.

Im Projektzeitraum der Fellow-Programm steht vor allem die Aufbereitung der bereits erschlossenen Daten zu den Losbüchern und Chiromantien für das FactGrid-Repository und die Entwicklung von Arbeitsroutinen zu Integration der Daten im Vordergrund. Über Vorträge, unter anderem im Netzwerk Historische Wissens und Gebrauchsliteratur soll das Projekt während dieser Zeit bekannt gemacht und mit einem Beitrag in der Open-Access-Zeitschrift „Mittelalter. Interdisziplinäre Forschung und Rezeptionsgeschichte“ vorläufig abgeschlossen werden.

Meilensteine

  1. Entwicklung eines Modells der Datenstruktur für die Einträge in FactGrid
  2. Entwicklung eines Routine zur Aufbereitung und Intergration der Daten
  3. Intergration der Daten zu den Chiromantien [student. Hilfskraft]
  4. Intergration der Daten der Dissertation (Das Losbuch. Manuskriptologie einer Textsorte des 14.-16. Jahrhundert. 2018) [student. Hilfskraft]
  5. Zeitschriftenbeitrag zur Überlieferung der deutschsprachigen Chiromantien bis ca. 1520
  6. Überarbeitung des Wikipedia-Artikels ‘Chiromantie’
  7. Zeitschrftenbeitrag über Projekt

Zwischenbericht Oktober/November 2020

Die Monate Oktober und November dienten vor allem zur organisatorischen Vorbereitung der Forschungsarbeit:

1. Personal:

Zunächst war unklar, ob und auf welche Wiese eine studentische Hilfskraft eingestellt werden kann. Dies ist nun über den Lehrstuhl für Ältere Deutsche Literatur der RWTH Aachen möglich. Der Abschluss eines Arbeitsvertrags ist aber erst dann sinnvoll, wenn die gesamt Summe des Stipendiums ausgezahlt wurde. Ansonsten wäre die Laufzeit zu kurz.

2. Forschungscommunity:

Aus einer seit 2019 bestehenden losen Forschergruppe heraus, haben wir am 4.10.2020 den Verein Netzwerk Historische Wissens- und Gebrauchsliteratur] (HWGL) gegründet und ich habe dessen Vorsitz übernommen. Der Verein organisiert regelmäßig Netzwerktreffen, in denen auch die kollaborative Datenspeicherung abgesprochen wird. Es besteht bereits ein von mir betriebenes Wiki. Ob sich FactGrid als Erweiterung desselben eignet, soll in diesem Projekt erkundet werden.

3. Datenmodell/Workshop:

Das Datenmodell soll gleichzeitig praktikabel und theoretisch fundiert sein. Es muss daher mit der Forschungscomunity abgestimmt werden. Einen ersten Entwurf werde ich am 10.12.2020 auf dem 1. HWGL-Abendkolloquium vorstellen. Auf einem Workshop im Rahmen des (virtuelles) Netzwerktreffen Historische Wissens- und Gebrauchsliteratur (08.-11.01.2021) soll der praktische Umgang mit diesem erprobt werden. Ich bin auch an der Organisation beider Veranstaltungen beteiligt. Bei der Gestaltung des Datenmodells sind auch rechtliche Aspekte wichtig. Ursprünglich war ein Import der gesamten Datenbestände des Handschriftencensus angedacht. Dies ist aus rechtlichen Gründen (Unterschiede in der Lizensierung), jedoch nicht möglich. Ein Import von Daten scheint mir nur dann rechtlich zulässig, wenn das Datenmodell der FactGrid Datenbank sich wesentlich von dem des Handschriftencensus unterscheidet.

4. Thematisch relevante Publikationen:

    • Zu den Artikel ‘Sortes’ und ‘German Texts on Superstition’ des Handbuchs Prognostication in the Medieval World mussten noch die Fahnen korrigiert werden. Das Handbuch ist am 09.11.2020 erschienen.
    • Das Manuskript eines Beitrags zur Analyse von handschriftlich Überliefertem mittels Graph-Datenbanken, habe ich im September abgeschlossen. Dieser Beitrag hat Einfluss auf das Datenmodell in FactGrid, denn die dort entwickelten Analyseprozesse sollen auch mit den späteren FactGrid-Daten möglich sein. Der Artikel wird in der Zeitschrift für digitale Geisteswissenschaften erscheinen und wird derzeit redaktionell eingearbeitet.
    • Gemeinsam mit Björn Reich und Matthias Standke gebe ich einen Sammelband mit Editionen früher gedruckter Losbücher heraus. Die eingereichten Texte werden von uns redigiert. Die Erschließungsdaten der Losbücher, die ich hauptsächlich bereits in meiner Dissertation erfasst habe, werden einer der ersten Datenbestände sein, die in FactGrid importiert werden.
    • Die Erschließungsarbeit für den Katalog der deutschsprachigen illustrierten Handschriften der Bayerischen Akademie der Wissenschaften läuft weiter. Auch die Dabei gewonnenen Daten sollen in FactGrid importiert werden.

5. Wissenschaftskommunikation:

Meine Mentorin Anita Runge hat mich durch ihre Perspektive auf meine Forschung davon Überzeugt, dass deren Gegenstände und Ergebnisse auch für ein breiteres Publikum interessant sind. Ich versuche mich daher stärker im Bereich Wissenschaftskommunikation zu engagieren.

    • Bereits seit Anfang des Jahres läuft die Zusammenarbeit mit dem Germanischen Nationalmuseum in Nürnberg zur Ausstellung Zeichen der Zukunft. Wahrsagen in Ostasien und Europa, die eigentlich Anfang Dezember 2020 eröffnet werden sollte. Für den Katalog zur Ausstellung habe ich drei Artikel verfasst.
    • Im Januar werde ich für eine Folge des Podcasts Anno PunktPunktPunkt zum Thema Zukunftsprognostik im 15. Jahrhundert zwischen Aberglaube, Wissenschaft und Spiel. Zur Beziehung zwischen Diskurswandel und Medienwandel interviewt. Die Folge soll im Februar 2021 erscheinen.
    • An der Universität Salzburg werde ich am 13.1.2021 einen Gastvortrag zum Thema Kristallsehen. Praktiken und Erklärungsmodelle von der Antike bis heute halten.

Bild:

Konrad Bollstatter: ‘Complexiones-Würfelbuch’
Augsburg 1455
Schreiber: Konrad Bollstatter, Illustrationen: Werkstatt Johannes Bämler
München, Bayerische Staatsbibliothek, Cgm 312, fol. 51v-52r
(Bild: Bayerische Staatsbibliothek, Montage: Marco Heiles, Lizenz: CC BY-NC-SA 4.0)