Are our Wikibase QueryServices about to mess up two millennia of historical dates?

It was in February 2019 at a conference dinner of medievalists in Jena when I was first confronted with the calendar problem which Wikibase had been posing ever since it had digested its first Julian calendar dates. I had given a Wikibase demonstration earlier that day and now I was sitting next to a medievalist who was ready to destroy me: “Wikibase”, he stated, “is a genuine disaster without anyone understanding it.”

I demanded to hear why that should be the case and the man asked me to show him just one medieval date from Wikidata. I had activated my phone and landed on biography c. 1500.

“See that small print?” he asked, “these dates are all noted as Gregorian before 1584.”

The qualifier was indeed peculiar. Why would they set a Gregorian date before 1582 and then mark it as such? “Well, you know, that the Gregorian calendar was only introduced in 1582, do you?!”

Of course I knew. I am an 18th-century person and Britain had introduced this calendar as late as 1752. The reform had by that time to close a gap of 11 days. But I could also point out that Wikibase allowed the fast correction: “You can easily switch between the calendars, and the machine will actually understand the implications on any timeline” I showed him my screen:

The man was in agony: “Too late. Wikidata is already in big shit”. I realised that I was lacking the full astronomical background and that I did not know the story of these peculiar Wikidata redactions.

Why we needed the Gregorian calendar in the first place

Both, the Julian calendar of 46 BC and the superior Gregorian calendar first introduced in 1582, are approximations. A solar year is one circle around the sun whilst the globe is spinning at about 365.2422 revolutions per year — year after year our planet ends its tour with a different slice pointing towards the sun. We are, in fact slowing down, thanks to the friction which the moon’s gravitation is generating in a constant movement of ebbs and tides, but that is another story. 365.2422 turns per year is our present spin more or less exactly but difficult to generate in a procedural long term pattern of constant adaptations.

The Julian calendar, as it was introduced under Julius Caesar in 46 BC, added one day every four years — in the so called leap years — a rule that boiled down to an additional quarter of a day per year. The approximation of 0.25 days against 0.2422 missed its mark just by 0.0078 days per year, less than a hundredth of a day — negligible one might think — but that one hundredth of a day is a day in a hundred years. In a millennium this discrepancy is piling up to 7.8 days, in two millennia to half a month, moving Christmas further and further away from the longest night until we can finally celebrate Christmas and Easter on the same day; and that was why the Gregorian calendar was finally introduced in 1582 with its far more complex regime of leap years:

  • add one day every four years (as you did under the Julian calendar to create a year of 365.25 days)
  • omit every leap year that is divisible by 100 to get a lower number
  • let this leap year, however, happen if it is divisibly by 400 in order to get a year of 365.2425 days.

The Gregorian calendar reduced the aberration to a surplus of 0.0003 days per year — it will now take 3333 years until we need an additional day to be back in tune with the solar year. The Vatican in Rome adopted the calendar on the 4th of October 1582 — jumping over night into Friday the 15th of that year. Christianity, however, was at that point no longer ready to obey a Papal decree. Eastern Orthodox churches stayed on the Julian calendar right into the the 20th century; Protestant territories and realms would decide one by one. Prussia (with its complex ties into catholic Poland adopted the new calendar in 1612 whilst most of the other Protestant territories stayed Julian for the next 88 years. The United Kingdom took the step in 1752. Lithuania, Russia, and Greece were to switch as late as 1915, 1918 and 1923 respectively.

The following map is from reddit:

When Europe switched from Julian to Gregorian calendar.
byu/coneyislandimgur ineurope

…and it is far from getting the full complexity. The following list gives the growing FactGrid table:

Europe was fragmented. Travelling across Germany in 1699, you could date your letters switching back and forth at every customs house on your tour:

Map of the Holy Roman Empire 1648. Wikimedia Commons

How we solved the problem — and created an even bigger one

Wikibase is a bright software. The tools — the QueryService and QuickStatements — are (or were) not immediately that bright, and that caused the mess the medievalist had noted. QuickStatemens, the tool for mass imports, simply did not offer a Julian calendar switch before February 2023. Instead it would mark all dates as Gregorian without asking — which, looking backwards, was not all that bad…

…why could we all live with the erroneous labelling of Julian dates as Gregorian on Wikidata? Because Wikidata was with this negligence basically doing what we all had been doing up to that point.

Johann Sebastian Bach was born on the 21st of March 1685. Germany’s central database, the GND, is stating this date up until now without the slightest remark on the calendar. The date is Julian because Eisenach’s church register was keeping records in the Julian calendar for another 15 years. The composer himself will not have shifted his birthday to the 31st of March in 1700, the year of the great reset. We all ignore the shift and keep copying dates from documents without any interference. Calendar experts might be interested in the “real” day and they can create Julian/Gregorian calendar matches in those rare cases in which they have to create an exact timeline of events with dates of both calendars.

The Wikidata community had been unaware of the problem. The Gregorian label on all the Julian days was foolish, but the input was actually stabilising our historical tradition as the QueryService will not do anything odd with dates that are entered as Gregorian.

I was far from seeing these advantages after my conversation of 2019 and that was why I warned the PhiloBiblon team in 2022 that QuickStatements would label all their Julian dates as Gregorian against all better intentions once they were imported to FactGrid. Charles Faulhaber immediately asked their programer, Josep Maria Formentí, whether he could not take a look into QuickStatements to solve that little problem. Weeks later Josep introduced the /J-switch that is now available to mark any date as Julian in mass inputs:

+ 1751-06-16T00:00:00Z/11/J

You can now feed thousands of medieval or early-18th-century British dates into your Wikibase and your machine will present all these dates in timelines in perfect synchrony with Gregorian dates. This is extremely nice if you are editing a correspondence whose partners were signing their letters under various calendars. Your machine will give you the exchange of letters in their true course.

So why the alarm?

Wikibase is an intelligent software; it brings objectivity into your statements. Feed a Julian date into your Wikibase and that day will be noted as Julian on the Wikibase itself.

Things get messy wherever we retrieve Julian dates from the QueryService, since this is where the production of funny (and eventually of erroneous) dates will be begin. The QueryService will convert all Julian dates into mathematically correct Gregorian dates.

Martin Luther is known to have died on the 18th of February 1546 — under the Julian calendar, that needs not to be stated, and our Wikibase is giving that date without any calendar stamp on it. But ask the QueryService for Luther’s birthday and it will tell you that the church reformer actually died on the 28th of March, 10 days later — a Gregorian calendar date (without indication) (no big issue you might think, now that you know).

And now think of masses of data which we will be moving between Wikibases in the brave new world of “federates Wikibases”. If there are “Julian” dates among them, then these will get secret additional days wherever they are extracted with the help of a regular SPARQL-Query on the QueryService.

This is what will happen to Luther’s date of death as it is now no longer a subject of safe copying. We will see it in an increasing number of variants — namely as:

  • 18 February 1546 (Greg.) — mistaken QuickStatements input artefact
  • 18 February 1546 (Jul.) — the historically correct date
  • 28 February 1546 (Greg.) — unorthodox but correct Wikibase QueryService output
  • 28 February 1546 (Jul.) — Wikibase output mistakenly saved as Julian
  • 10 March 1546 — the previous converted to Gregorian
  • 20 March 1546 — the previous after the next im- and export

and so on and so on.

Can we stop the wave of uncontrolled additions of days on Julian calendar dates?

I am not quite sure how. We need a QueryService that will never ever offer a historical date without the corresponding calendar statement (now that we have a machine that does both calendars).

But not only the QueryService is posing a problem here. Our Wikibases should have a third option, because our documentary evidence is usually lacking calendar information. Eisenach’s church register of 1685 is using the Julian Calendar (without further notice), that is something we can determine — but we cannot say what calendar an author of a typical letter was using in 1685 if that date comes without a localisation. Our documents do not tend to have calendar statements on them.

What we need here is a third — a “calendar format unknown” — option. It’s complicated, I am afraid.

Links and more

  • Header image from Ολυμπία δώματα, or, An almanack for the year of our Lord God 1752 (London: Printed by T. Parker, for the Company of Stationers, 1752), from the digitisation at Archive.org
  • English Wikipedia List of adoption dates of the Gregorian calendar by country https://en.wikipedia.org/
  • See also: Maniphest T207705, Implement the Extended Date/Time Format Specification, https://phabricator.wikimedia.org/T207705
  • Lydia Pintscher, calendar model screwup, 30 Jun 2015. [https://lists.wikimedia.org/hyperkitty/list/wikidata@lists.wikimedia.org/thread/Y7OEHUYV66DHRVZ6JCSODWAYZ25SLUHM/ https://lists.wikimedia.org/hyperkitty]
  • Julian and Gregorian dates from Wikidata, question asked on https://opendata.stackexchange.com/, Apr 18, 2018 at 0:33 [https://opendata.stackexchange.com/questions/12723/julian-and-gregorian-dates-from-wikidata https://opendata.stackexchange.com/]

PhiloBiblon: From Siloed Databases to Linked Open Data via Wikibase: Proof of Concept

We are very pleased to announce a pilot project funded by the U.S. federal government’s National Endowment of the Humanities (NEH): “PhiloBiblon: From Siloed Databases to Linked Open Data via Wikibase: Proof of Concept.” It will begin June 1, 2021. and end May 30, 2022 and will be hosted by FactGrid.

The project is focused on PhiloBiblon. a forty-year-old database for the study of the medieval history and literatures of the Romance cultures of the Iberian Peninsula: Portuguese and Galician-Portuguese, Castilian, and Catalan. It contains four subsidiary databases:

  • BETA: Bibliografía Española de Textos Antiguos: Medieval texts in Spanish.
  • BIPA: Bibliografía de la Poesía Áurea: Golden Age (16th-17th c.) Poetry in Spanish.
  • BITAGAP: Bibliografia de Textos Antigos Galegos e Portugueses: Medieval texts in Galician, Galician-Portuguese, and Portuguese,
  • BITECA: Bibliografia de Textos Antics Catalans, Valencians i Balears: Medieval texts in Catalan.

The project is designed to solve one of the most vexing problems facing long-standing digital projects: maintainance of the software platform. PhiloBiblon started out in 1975 as an ancillary database of the Dictionary of the Old Spanish Language project, carried out at University of Wisconsin, Madison, by Lloyd Kasten and his student, John Nitti. In Madison the database management system (DBMS) used was FAMULUS, created, ironically, in Berkeley in 1964, for the Pacific Southwest Forest and Range Experiment Station. Since then, technological transformation has been a constant: from the CD-ROM discs of ADMYTE (Archivo Digital de Manuscritos y Texts Españoles ) to a first web version in 1997 to the current 2014 version. Each transformation has usually required multiple grant applications to agencies and foundations, especially to NEH.

PhiloBiblon currently runs under Windows on Revelation Technology’s MultiValue database, OpenInsight. PhiloBiblon’s 1987 implementation was designed on Revelation G., the ancestor of OpenInsight, by John May, a graduate student in History of Science who eventually left the academy for a business career. John has maintained and enhanced the PhiloBiblon DBMS for more than 35 years, building it out to encompass ten relational tables (texts, witnesses, primary source manuscripts and imprints, copies of imprints, persons, institutions, toponyms, and secondary references). Among them these ten tables contain 1246 data elements (fields), 98 controlled vocabulary lists with more than 3000 properties, 110 search indexes, and 30 data entry screens.

PhiloBiblon and its relational DBMS existed before Tim Berners-Lee invented the Worldwide Web at CERN in Geneva in 1989. It was evident from the advent of the first really useful commercial browser, Netscape, that the WWW offered a vastly superior vehicle for making information available to students and scholars, much better than print and CD-ROM. In PhiloBiblon’s first web version (1997), data were exported from the OpenInsight DBMS and uploaded to a web server at Berkeley in HTML format. Users could search only for authors, titles, or keywords; and the result of the search was just a list of manuscripts or editions, each of which had to be opened, one-by-one, to find the item of interest.

In the current version OpenInsight data is still exported—usually every two or three months—and uploaded to a web server at Berkeley (and to a mirror site at the Universitat Pompeu Fabra in Barcelona) in ten separate files, one for each table, in the more advanced XML format. On the server the eXtensible Text Framework (XTF) program, created by the California Digital Library of the University of California, parses each file into individual records, indexes them, and serves them to users as a result of a search request from the PhiloBiblon web site. The search mechanisms are much more powerful, offering not only keyword searching but also a set of search boxes tailored to each entity.

Thus for texts (works): keyword (simple search), author, title, incipit, explicit, associate person, date and place of composition, subject.

For manuscripts and editions: City and library holding the item, shelfmark, date and place of production, printer and publisher, scribe and patron, previous owner or other associated person:

The time lag between data input and its appearance on the web is annoying. Moreover, the process required to export data and upload it to the web to the web is is neither elegant nor efficient. Aside from this and from the problem of maintaining and enhancing both the Windows DBMS and the web software, the most urgent issue facing PhiloBiblon is its status as an information silo, with no organic relationship with other information sources. The web 3.0, the semantic web, is designed to make use of Linked Open Data and the Resource Description Framework (LD/RDF) in order to make possible automatic links to other information resources, like the Virtual International Authority File (VIAF).

Since 2014 the PhiloBiblon research teams have prepared a series of unsuccessful grant proposals, separately as well as in collaboration with other projects, to NEH and to Spanish, Catalan, and European agencies and foundations with the goal of funding the transformation of PhiloBiblon into an LD/RDF resource.

Last year, instead of proposing the creation ex professo of a new web-based DBMS for PhiloBiblon, on the advice of our neighbors at Stanford University, the University of California, Davis, and the international library consortium OCLC, we decided to explore a radically different solution: to incorporate PhiloBiblon into the wiki world. PhiloBiblon already cites Wikipedia constantly, more than 1400 times in nine different languages. Of more interest than Wikipedia as a model, however, is the more structured but still open environment of Wikidata. Wikidata, however, is too open. PhiloBiblon requires more control over the individuals who can contribute to it.

This led us to FactGrid, which is ideal for our purposes. Open only to members, it offers a perfect sandbox for the PhiloBiblon staff to explore the relationship between PhiloBiblon’s highly structured data model and the elegant and infinitely extensible model of triplestores based on Q# entities and P# properties. We are enormously grateful to Olaf Simons not only for his generous offer to make it available for this purpose but also for agreeing to serve on the Advisory Board of the current NEH project.

When this project ends May 31, 2022, we hope to have shown that the Wikidata model is viable for PhiloBiblon over the long term and that we can make use of its standard input processes, modified as necessary, to map PhiloBiblon’s 421,000 records into the corresponding FactGrid entities and properties, creating new ones as necessary. This work will be carried out primarily by data analyst Adam Anderson, whose academic specialization is Assyriology and cuneiform studies. He will also study Wikibase’s LD access points to and from libraries and archives and test the Wikibase data export module for JSON-LD, RDF, and XML on PhiloBiblon data.

TABLA BETA BITAGAP BITECA BIPA
ANALYTIC (witnesses) 14692 52084 12239 89926
REFERENCES 7270 21558 5976 472
PERSONS 7423 32309 3473 3418
GEOGRAPHY 1814 4759 840 211
INSTITUTIONS 794 3297 585 4
LIBRARIES 915 455 420 119
MANUSCRIPTS & IMPRINTS 5168 5886 1971 1572
COPIES OF PRINTED BOOKS 4157 1146 1473 137
SUBJECT HEADINGS 339 34 149 126
WORKS (texts) 6034 31962 6173 89913
TOTAL 48606 153490 33299 185898 421293

In addition, software engineer Josep María Formentí (Barcelona), after evaluating the Wikibase data entry module and report format, will create prototypes of more user-friendly query and data entry screens and report formats.

All of this work will be carried out in collaboration with the twenty members of the PhiloBiblon volunteer academic staff and, we hope, numerous volunteers from the Hispano-medievalist community.


Image: Rueland Frueauf the Elder (1440–1507) The Education of the Infant Christ (1506) Wikimedia Commons.