PhiloBiblon receives a new grant from the National Endowment for the Humanities

We are delighted to announce that PhiloBiblon, a database of the primary sources for the study of medieval Iberia,  has received a two-year implementation grant from the Humanities Collections and Reference Resources program of the National Endowment for the Humanities to complete the mapping of PhiloBiblon from its almost forty-year-old relational database technology to the Wikibase technology that underlies Wikipedia, Wikidata, and FactGrid. The project will start on the first of July and, Dios mediante, will finish successfully by the end of June 2025.

The fundamental problem is to map the 422,000+ records of PhiloBiblon’s bibliographies with their complexly interrelated relational tables to the triplestore structure of Wikibase. A triplestore relates two Items by means of a Property. Thus a Work is linked to an Author by the Property “written by.”

We received an NEH Foundations grant for this project in 2021, as described in detail in PhiloBiblon: From Siloed Databases to Linked Open Data via Wikibase: Proof of Concept. Over the course of the last two years, the pilot project team, consisting of Charles Faulhaber (PI), Patricia García Sánchez Migallón and Almudena Izquierda (doctores por la UCM), Berkeley undergraduate Spanish and data science majors (Julieta Soto, Serena Bai, Tina Lin, Cassandra Calciano, Martín García Ángel), Max Ziff (data engineer), and Josep Formentí (user interface programmer), with the guidance of Olaf Simons, has analyzed the data structures of PhiloBiblon’s ten relational tables (using BETA for the test cases) and worked out the procedures needed to convert them into triplestore structures.

Almudena and Patricia manually mapped more than 125 BETA records to FactGrid: PhiloBiblon as models for the automated processing of the rest. See for example the records for Alfonso X, BNE MSS/10069 (Cantigas de Santa Maria), and the 1497 edition of the translation of Boccacio’s Fiammeta. These models have been key for establishing the semantic relations between PhiloBiblon’s data fields and the Properties and Items in FactGrid. In many cases appropriate properties did not exist and it was necessary to create them. For example, something as simple as the Watermark property was needed in order to identify the various watermark types set forth in PhiloBiblon’s controlled vocabulary.

Julieta Soto and Martín García Ángel attacked the problem of creating almost 900 FactGrid records for the controlled vocabulary terms in BETA. This meant in the first place a search in FactGrid to make sure that an equivalent term did not already exist, in order to avoid creating duplicate records. Then they had to situate the term in the FactGrid ontology by specifying it as a “basic object” (e.g., fruit) or identifying it as a subclass of an appropriate basic object, for example facsímil impreso as a subclass of facsímil. At the same time they had to link the record to the code in PhiloBiblon, BIBLIOGRAPHY*RELATED_BIBCLASS*FAP, identifying a record in the Bibliography table as a print facsimile, thereby making it possible to search for such items.

The default viewer used in FactGrid, the same as that used in Wikidata, is not user friendly. Therefore Josep has created a prototype user interface, using data from the BETA Institutions table. We encourage you to play with it and tell us what you like or—more usefully—don’t like.

This change to Wikibase technology is designed to allow PhiloBiblon not only to take advantage of the linked open data of the semantic web but also, and most importantly, to decrease sustainability costs. Because Wikibase is open-source software maintained by WikiMedia Deutschland, the software development arm of the Wikimedia Foundation, software maintenance costs for PhiloBiblon will be minimal in the future. This means that it will no longer be necessary to seek major grant support every five to seven years merely to keep up with technology change.

While this work has been going on, we have not neglected the vital process of cleaning up PhiloBiblon data in order to facilitate the automated mapping nor the equally vital process of adding new information to PhiloBiblon. For example, Pedro Pinto, a member of the BITAGAP team, has recently discovered a “folha desmembrada” (BITAGAP manid 7862) from the Livro 4 of the chancery records of king Fernando I (1345-1383) (BITAGAP manid 3255), separated from the manuscript in the Arquivo Nacional da Torre do Tombo. The newly discovered dismembered leaf contains five previously unkown royal documents. It was being used as the cover of the “Livro de Acordãos, 1620-24,” in the archive of the Santa Casa de Misericórdia in Coruche, a small city in the Santarem district on the Tagus river northeast of Lisbon.

The recycling of parchment leaves from discarded medieval manuscripts, presumably for more socially beneficial purposes, such as the protection of administrative records, was common in both Spain and Portugal in the sixteenth and seventeenh centuries. Such leaves have been the source of many previously unrecorded medieval texts. Perhaps the most spectacular exemple was Harvey Sharrer’s discovery in 1990 of the eponymous Pergaminho Sharrer (BITAGAP manid 1817), with seven unknown poems of king Dinis of Portugal (1279-1325). This had been used as the binding of a collection of notarial documents (Lisboa: Arquivo Nacional da Torre do Tombo: Lisboa, Cartório Notarial de. N. 7-A, Caixa 1, Maça 1, livro 3).

PhiloBiblon: From Siloed Databases to Linked Open Data via Wikibase: Proof of Concept

We are very pleased to announce a pilot project funded by the U.S. federal government’s National Endowment of the Humanities (NEH): “PhiloBiblon: From Siloed Databases to Linked Open Data via Wikibase: Proof of Concept.” It will begin June 1, 2021. and end May 30, 2022 and will be hosted by FactGrid.

The project is focused on PhiloBiblon. a forty-year-old database for the study of the medieval history and literatures of the Romance cultures of the Iberian Peninsula: Portuguese and Galician-Portuguese, Castilian, and Catalan. It contains four subsidiary databases:

  • BETA: Bibliografía Española de Textos Antiguos: Medieval texts in Spanish.
  • BIPA: Bibliografía de la Poesía Áurea: Golden Age (16th-17th c.) Poetry in Spanish.
  • BITAGAP: Bibliografia de Textos Antigos Galegos e Portugueses: Medieval texts in Galician, Galician-Portuguese, and Portuguese,
  • BITECA: Bibliografia de Textos Antics Catalans, Valencians i Balears: Medieval texts in Catalan.

The project is designed to solve one of the most vexing problems facing long-standing digital projects: maintainance of the software platform. PhiloBiblon started out in 1975 as an ancillary database of the Dictionary of the Old Spanish Language project, carried out at University of Wisconsin, Madison, by Lloyd Kasten and his student, John Nitti. In Madison the database management system (DBMS) used was FAMULUS, created, ironically, in Berkeley in 1964, for the Pacific Southwest Forest and Range Experiment Station. Since then, technological transformation has been a constant: from the CD-ROM discs of ADMYTE (Archivo Digital de Manuscritos y Texts Españoles ) to a first web version in 1997 to the current 2014 version. Each transformation has usually required multiple grant applications to agencies and foundations, especially to NEH.

PhiloBiblon currently runs under Windows on Revelation Technology’s MultiValue database, OpenInsight. PhiloBiblon’s 1987 implementation was designed on Revelation G., the ancestor of OpenInsight, by John May, a graduate student in History of Science who eventually left the academy for a business career. John has maintained and enhanced the PhiloBiblon DBMS for more than 35 years, building it out to encompass ten relational tables (texts, witnesses, primary source manuscripts and imprints, copies of imprints, persons, institutions, toponyms, and secondary references). Among them these ten tables contain 1246 data elements (fields), 98 controlled vocabulary lists with more than 3000 properties, 110 search indexes, and 30 data entry screens.

PhiloBiblon and its relational DBMS existed before Tim Berners-Lee invented the Worldwide Web at CERN in Geneva in 1989. It was evident from the advent of the first really useful commercial browser, Netscape, that the WWW offered a vastly superior vehicle for making information available to students and scholars, much better than print and CD-ROM. In PhiloBiblon’s first web version (1997), data were exported from the OpenInsight DBMS and uploaded to a web server at Berkeley in HTML format. Users could search only for authors, titles, or keywords; and the result of the search was just a list of manuscripts or editions, each of which had to be opened, one-by-one, to find the item of interest.

In the current version OpenInsight data is still exported—usually every two or three months—and uploaded to a web server at Berkeley (and to a mirror site at the Universitat Pompeu Fabra in Barcelona) in ten separate files, one for each table, in the more advanced XML format. On the server the eXtensible Text Framework (XTF) program, created by the California Digital Library of the University of California, parses each file into individual records, indexes them, and serves them to users as a result of a search request from the PhiloBiblon web site. The search mechanisms are much more powerful, offering not only keyword searching but also a set of search boxes tailored to each entity.

Thus for texts (works): keyword (simple search), author, title, incipit, explicit, associate person, date and place of composition, subject.

For manuscripts and editions: City and library holding the item, shelfmark, date and place of production, printer and publisher, scribe and patron, previous owner or other associated person:

The time lag between data input and its appearance on the web is annoying. Moreover, the process required to export data and upload it to the web to the web is is neither elegant nor efficient. Aside from this and from the problem of maintaining and enhancing both the Windows DBMS and the web software, the most urgent issue facing PhiloBiblon is its status as an information silo, with no organic relationship with other information sources. The web 3.0, the semantic web, is designed to make use of Linked Open Data and the Resource Description Framework (LD/RDF) in order to make possible automatic links to other information resources, like the Virtual International Authority File (VIAF).

Since 2014 the PhiloBiblon research teams have prepared a series of unsuccessful grant proposals, separately as well as in collaboration with other projects, to NEH and to Spanish, Catalan, and European agencies and foundations with the goal of funding the transformation of PhiloBiblon into an LD/RDF resource.

Last year, instead of proposing the creation ex professo of a new web-based DBMS for PhiloBiblon, on the advice of our neighbors at Stanford University, the University of California, Davis, and the international library consortium OCLC, we decided to explore a radically different solution: to incorporate PhiloBiblon into the wiki world. PhiloBiblon already cites Wikipedia constantly, more than 1400 times in nine different languages. Of more interest than Wikipedia as a model, however, is the more structured but still open environment of Wikidata. Wikidata, however, is too open. PhiloBiblon requires more control over the individuals who can contribute to it.

This led us to FactGrid, which is ideal for our purposes. Open only to members, it offers a perfect sandbox for the PhiloBiblon staff to explore the relationship between PhiloBiblon’s highly structured data model and the elegant and infinitely extensible model of triplestores based on Q# entities and P# properties. We are enormously grateful to Olaf Simons not only for his generous offer to make it available for this purpose but also for agreeing to serve on the Advisory Board of the current NEH project.

When this project ends May 31, 2022, we hope to have shown that the Wikidata model is viable for PhiloBiblon over the long term and that we can make use of its standard input processes, modified as necessary, to map PhiloBiblon’s 421,000 records into the corresponding FactGrid entities and properties, creating new ones as necessary. This work will be carried out primarily by data analyst Adam Anderson, whose academic specialization is Assyriology and cuneiform studies. He will also study Wikibase’s LD access points to and from libraries and archives and test the Wikibase data export module for JSON-LD, RDF, and XML on PhiloBiblon data.

TABLA BETA BITAGAP BITECA BIPA
ANALYTIC (witnesses) 14692 52084 12239 89926
REFERENCES 7270 21558 5976 472
PERSONS 7423 32309 3473 3418
GEOGRAPHY 1814 4759 840 211
INSTITUTIONS 794 3297 585 4
LIBRARIES 915 455 420 119
MANUSCRIPTS & IMPRINTS 5168 5886 1971 1572
COPIES OF PRINTED BOOKS 4157 1146 1473 137
SUBJECT HEADINGS 339 34 149 126
WORKS (texts) 6034 31962 6173 89913
TOTAL 48606 153490 33299 185898 421293

In addition, software engineer Josep María Formentí (Barcelona), after evaluating the Wikibase data entry module and report format, will create prototypes of more user-friendly query and data entry screens and report formats.

All of this work will be carried out in collaboration with the twenty members of the PhiloBiblon volunteer academic staff and, we hope, numerous volunteers from the Hispano-medievalist community.


Image: Rueland Frueauf the Elder (1440–1507) The Education of the Infant Christ (1506) Wikimedia Commons.