A FactGrid vocabulary of types of functional texts after Eckard Rolf

auf Deutsch

The data set in basic queries:

The links above give access to our first controlled vocabulary on FactGrid: “The FactGrid vocabulary of types of functional texts.”

Types of functional texts are not a matter of course. The corresponding genres of literary texts are, with all the problems discussed in the literary debate, far better known. The genres of functional texts are less controversial, but also less comprehensive. You find them in any mass of public records in the form of “insurance policies,” “interrogation protocols,” and “school reports,” to the odd “delousing certificate.” They do their jobs – so why collect the terms?

The technical answer is that Wikibase is software that invites you to use very specific vocabularies. You can run SPARQL queries on these specific terms and they will retrieve the “delousing certificate” in the mass of data, and you can, with very simple switches, bring far broader fields into view. The reduction to fields is the first step into statistics. If you can bring the variety under broader headings you can get a quick view of any vast production in your table. The vocabularies you need for this purpose have to be more than just lists of words. They need categorisations, common denominators above the words, ontologies.

A linguist’s perspective and data model

Our “Vocabulary of types of functional texts” has the required structural depth to allow statistical analysis and the broader analysis of larger bodies of texts. So far it is based primarily on the two Properties P894 “Eckard Rolf class of functional text types” and P912 “Speech act qualities.” The analysis is under both properties based on Eckard Rolf’s Die Funktionen der Gebrauchstextsorten (Berlin/ New York, 1993). Tobias Christ asked for the import of this vocabulary and its inherent structure for a project on functional texts of Germany’s Nazi era. He will explore handbooks for the organisers of Hitler Youth camps, official directives on the insignia of uniforms, etc. Rolf’s book is immensely practical with its in-depth analysis of 2055 terms arranged here in the five branches of illocutionary acts. Types of functional texts, under this premise, are essentially illocutionary speech acts as proposed by Austin and Searle in the 1950s and 1960 in their five branches:

  • assertives = speech acts that commit a speaker to the truth of the expressed proposition,
  • directives = speech acts that are to cause the hearer to take a particular action, e.g. requests, commands and advice,
  • commissives = speech acts that commit a speaker to some future action, e.g. promises and oaths,
  • expressives = speech acts that express on the speaker’s attitudes and emotions towards the proposition, e.g. congratulations, excuses and thanks
  • declarations = speech acts that change the reality in accord with the proposition of the declaration, e.g. baptisms, pronouncing someone guilty or pronouncing someone husband and wife

…thus the Wikipedia article illocutionary acts. Rolf deployed three further layers underneath this basic differentiation: two layers (of more or less specific) options on how the respective aims can be achieved and the fourth layer of situational conditions. The “delousing certificate” is under this matrix a “declarative statement” (a person is “declared” to be free of lice after the required treatment). The certificate will add a “personal dimension” to the bearer of the certificate – he or she will be free again to interact with others with the legitimation of the certificate. The statement is finally “body related.” Rolf created 100 groups under these four layers. The “delousing certificate” is in group “DECLA 12” together with the “allergy passport” or the “vaccination certificate.” Other types of texts do different things differently: A “doctoral thesis” (ASS 24) is an “assertive” – it commits the speaker to the truth of his or her exploration. The work is supposed to be “descriptive” and “argumentative.” The author will hand in this work with the “intention to gain a specific qualification.” Neighbouring types of texts such as the “book review” share some but not all features: Book reviews are again “assertives” and “descriptive” but without the author’s intention to gain a specific qualification with them. Their focus lies on a “judgment” they pass.

The following search gives the entire vocabulary in the four languages that are presently fully supported with Eckard Rolf’s primary classes and their basic categorisation:

  • Types of functional texts, generic terms in German, English, French, Spanish with Eckard Rolf’s bottom-line classification https://tinyurl.com/27ctrfam

The actual set of words is – especially on its German side – larger than the set of items. Rolf had separated terms like “Jagdschein” and “Jagdkarte” – in this case to have the German and the Austrian terms. In English both things are “hunting permits” unless we decide to offer individual  hunting permits all around the world. In other cases the differences were stylistic, created by registers that could not be reproduced in English, French, or Spanish. We eventually reduced the set to objects of essentially the same meaning. About 200 words are now variants in the alias sections and on the P34 “naming” Property where they can attract explanations of their proper use.

Eckard Rolf presented his original 2055 terms together with a series of structural overviews. Tobias Christ offered a condensed PDF-version of these visualisations in a single tree structure. You can use the EntiTree App to show rhis structure on your screen:

Any of the nodes in the representation can determine a specific query of the terms at the end of the ensuing ramification. This is the complete list of speech act qualities searchable on the P912 Property of “Qualities of speech acts”:

It is just as easy to generate statistics on each structural level. Here is the visualisation of the top level in a bar chart:

The following four searches give the scripts for each level (the level difference is determined with the Q-Item in line 8):

  1. The five illocutionary purposes of speech acts Q538467
  2. Division by general way to achieve the purpose Q538468
  3. Division by specific way to achieve the purpose Q538469
  4. Division by primary conditions of speech acts Q538470

The searches above cover the entire terminological set so they can now be run on any specific body of texts.

The open tool

The FactGrid database version of Eckard Rolf’s structural analysis should turn the book’s considerations into an immensely practical tool ready to download into any other software environment and ready to be expanded on FactGrid. New types will not compromise the original set – it remains intact through the statement P124+Q514322 (“listed in Eckard Rolf, Die Funktionen der Gebrauchstextsorten”) that is made on every individual word:

The best way to add a new term is to find neighbouring terms and to adopt their statements. FactGrid already had a couple of candidates, such as the popular “Briefsteller” (the “letter writer’s guide”) of the German 18th century, or the “Quibus Licet” (the letter which Illuminati had to hand in every month to stay in contact with the “unknown superiors”).

Our first “controlled vocabulary” is with these preliminary remarks still very much of an experiment. —

  • It will be interesting to offer the generic terms also as “Lexemes”: — Wikibase Lexemes are special entities that organise individual words in their languages.
  • The French and Spanish labels in particular are still very artificial translations – we should have original terms for each of these items as referenced in historical documents.
  • The present vocabulary is not yet matched with external databases. Wikidata and the GND are the two most urgent data partners here. The following search gives the matching so far: https://tinyurl.com/284jbjhc
  • The linguist’s categorisation should be seen as one option to make sense of all these terms. One can easily think of other qualities of speech acts. In our preliminary talk Rolf proposed to explore, for instance, the assumed-sincerity dimension in many of these speech acts. “Lip service” was his example – a speech act where the “honesty of the emittent is unclear or doubtful.” One can just as well create completely independent properties on features of genres beyond the linguist’s interest.
  • We should eventually expand this work. A vocabulary of genres in all the arts and literature would be of interest here. One would balance such a vocabulary with a particular vocabulary of “historical generic terms.” (I remember, I once wrote a 700 page book with a plea to explore these terminologies in all their “deficiencies.” The deficiencies, so I proposed back then, were usually the first indications that people were not doing the things we are doing with works of “art” and “literature” in our debates. We might question our keenness on succinct definitions in these particular fields, so my thought ages ago; we do not really define words in order to settle debates, we are always far more interested in the destabilisation the definition will actually produce – but that is already a topic for a very different blog post.)

Published as part of the NFDI4Memory Task Area “Data Connectivity”, Historical Data Center Halle, project number 501609550.

9 x FactGrid, Coffee Talk Serie an der Universität Erfurt, 14. April – 16. Juni 2022, Donnerstags 13:30

Die Universität Erfurt lud uns ein, im kommenden Sommersemester eine online Coffee-Talk Serie zum FactGrid als kollaborativer Forschungsplattform zu veranstalten. Neun Themenschwerpunkte haben wir ausgesucht. Die Veranstaltungen sollen kurz und für die Mittagspause zum Hineinschnuppern gemacht sein. Lassen Sie sich inspirieren. Wir bieten eine 15minütige Erkundung mit jeweils offener Fragerunde.

Das Link zur Veranstaltung erhalten Sie für eine Mail an olaf.simons@pierre-marteau.com


14.4.2022: Sich beim Forschen über die Schulter sehen lassen? — Isabella Schwaderers Erkundungen zu den Mitgliedern der Schopenhauergesellschaft

Netzwerkverbindungen in der Schopenhauergesellschaft, 1912, 1913

Kann man es riskieren, auf einer Plattform, auf der alle Daten unmittelbar offen zugänglich sind, die eigene gerade erst angefangene Forschung laufen zu lassen? Isabella Schwaderer tat diesen Schritt mit ihren Recherchen zu den Mitgliedern der Deutschen Schopenhauergesellschaft und wird hier Einblicke in die Nutzerperspektive geben. Worauf lässt man sich ein? Was ist praktisch? Was ist unpraktisch? Was riskiert man? Was gewinnt man?


21.4.2022: Was immer eine Aussage finden kann, kann ein Datenbank-Item werden — wie Wikibase funktioniert

Wikibase steht im Ruf, ganz beliebige Information aufnehmen zu können. Auf einer einzigen Instanz kann man Information ganz verschiedener Fächer zusammenlaufen lassen und sie nahtlos über alle Fachgrenzen hinweg durchdringen.

Das Geheimnis liegt in der Flexibilität Tripel-basierter Daten. Wir können beliebige Objekte aufmachen und Aussagen zu ihnen beliebig an dokumentierte Datenstrukturen anpassen. Die Eingabe ist einfach. Komplizierter und offener ist, wie man die Daten danach in ihrer ganzen Vernetzung intelligent auswertet.

Ein Blick in die Datenmodellierung, die FactGrid Sample Searches und den Query-Service.


28.4.2022: Daten in 400 Sprachen verfügbar machen

Jahrzehntelang kämpfte man in der Bibliothekslandschaft um globale Datenmodelle und verbindliche Datenbank-Feldbelegungen in der Hoffnung, auf sichere Standards. All das hat das Wikidata-Projekt in seiner mutmaßlichen Notwendigkeit relativiert mit dem Angebot einer einzigen Ressource, die jede in ihr gespeicherte Aussage jederzeit in über 400 Sprachen verfügbar macht.

Wie das geht, ist im wörtlichen Sinne trivial: Alle Aussagen werden zerlegt in Datentripel von jeweils zwei Objekten und einer Beziehung zwischen ihnen, deren Teile man nun einzeln in beliebigen Sprachen mit beliebig vielen Labeln belegen kann. Tatsächlich können auf einer solchen Plattform Nutzer, ohne noch über eine gemeinsame Sprache zu verfügen, die Daten aller anderen in der eigenen Sprache lesen – eine gewaltige Chance für Projekte, die in Teams über Sprachgrenzen hinweg zusammenarbeiten sollen.

Ein Blick auf das Wikidata-Projekt, seine Software und Mehrsprachigkeit im FactGrid.


5.5.2022: Selbstorganisation über Projektgrenzen hinweg

Das FactGrid arbeitet ohne zentrale Redaktion, die Daten erst einmal ansehen und auf ihre Qualität hin überprüfen würde. Auch gibt es keine “Relevanzkriterien” – keine Kriterien, die festlegen, was für Daten in die Datenbank dürfen. Wir arbeiten mit einer verwirrenden Offenheit, die dafür ganz eigene Grenzen hat: Forschungsprojekte (auch private) stehen für ihre Arbeit extrem transparent ein. Wir sind hier an einigen interessanten Stellen anders organisiert als Wikidata.

Wie das in der Praxis geht, welche Konflikte man einkalkulieren und welche Konfliktzonen man eher nicht fürchten sollte – Erkundungen der Plattform-Architektur und der speziellen Freiräume, die wir in ihr Projekten gewähren.


12.5.2022, unusual time 18:00 CET: 350,000 objects with cuneiform inscriptions or: Data as a universal language — session in English with Adam Anderson, Berkeley

This is perhaps the most exciting FactGrid project at the moment – designed to create and to interconnect objects for all 350,000 cuneiform artifacts that known today. Where were these objects found? What events, what people, what places are noted on these objects? As a Wikibase installation we would serve as a database a wide collectively could work on and edit simultaneously in all its various present languages. Data stored on the platform would link into other databases and they would be uniquely easy to download for further work in all other software environments.

Our discussion was about the sheer quantities of data such a platform might eventually handle – if we went into the very texture of these objects, locating not only pieces of information but in further steps all the characters on all these objects in 3D data of the artifacts themselves.


19.5.2022: Georeferenzierte Objekte — Session with Bruno Belhoste in English

Paris to download (click on the map) / link for the direct table download TSV formatted

Eine Aufgabe, vor der DH-Projekte immer wieder stehen, ist es, Information auf Landkarten zu visualisieren. Räumliche Beziehungsnetze werden sichtbar, Nähe wird greifbar wie der Horizont, den Verfasser mit Korrespondenzen hatten. Soziale Phänomene, etwa die Zusammensetzung der Bevölkerung in verschiedenen Stadtvierteln, lassen sich erfassen. FactGrid-Information ist jederzeit georeferenzierbar. Wir bieten Georeferenzierungen in einem ersten Projekt – Paris to Download – zur beliebigen Nutzung auf der Plattform oder in anderen Software-Umgebungen an. Einige Blicke auf die Projekte und die Software, die hier nach neuen Modulen ruft.


9.6.2022: Genealogie im FactGrid – mehr als nur Väter und Mütter

What came after Robinson Crusoe’s first edition? EntiTree Visualisation

FactGrid-Information lässt sich komplex in externe Projekte hineinspielen. Zwei FactGrid Browsing-Tools stehen zu Verfügung. Es lassen sich jedoch auch ganz andere Werkzeuge denken.

Als überraschend vielseitig verwendbar erweist sich die von Orlando Groppo und Martin Schibel entwickelte EntiTree-App, die Genealogien auf bequeme Art und Weise mehrsprachig sichtbar macht. Spannende ist dass sich mit der EntiTree App auch noch ganz andere genealogische Beziehungen darstellen lassen.


16.6.2022: Die Zukunft im NFDI4Memory Gefüge – oder: Dateninseln zu neuem Leben erwecken

Das FactGrid ist seit 2021 gesetzt, um im geplanten NFDI4Memory-Konsortium der deutschen Geschichtswissenschaften als Wikibase-Instanz zur Verfügung zu stehen. Forschungsdatenmanagement ist hier das Thema. Was geschieht mit Forschungsdaten, die am Ende irgendwie übrigbleiben – gesammelt, um den Arbeitsprozess zu begleiten, doch danach irgendwie nutzlos, indes voller Korrekturen und Einblicke, die zukünftiger Forschung nutzen sollten? Was geschieht mit Daten, die bislang auf einer eigenen Plattform laufen, nachdem deren Förderung endet? Wie kann man Daten langfristig sichern?

Das FactGrid will hier die Ressource sein, die Information kollektiv nutzbar macht und langfristig in Zirkulation und Korrektur hält. Praktische Tipps, wie das gehen könnte.