Collaborating on the sum of all knowledge across languages

[The following article is from the Wikipedia @ 20 blog and extremely interesting to have here as well. The present copy is a draft version and offered on the original page together with the invitation to propose improvements. Denny Vrandečić is working at Google. Previously he has been at the Institute AIFB at the KIT (Karlsruhe Institue of Technology) and at Wikimedia Deutschland, where he founded Wikidata.]

Wikipedia is available in almost 300 languages, each with independently developed content and perspectives. Sharing more knowledge across languages would allow each edition to focus on their unique contributions, and yet improve their comprehensiveness and currency.

Differences between Wikipedia language editions

Wikipedia is often described as a wonder of the modern age. There are more than 50 million articles in almost 300 languages. The goal of allowing everyone to share in the sum of all knowledge is achieved, right?

Not yet.

The knowledge in Wikipedia is unevenly distributed. Let’s take a look at where the first twenty years of editing Wikipedia have taken us.

The number of articles varies between the different language editions of Wikipedia: English, the largest edition, has more than 5.8 million articles, Cebuano — a language spoken in the Philippines — has 5.3 million articles, Swedish has 3.7 million articles, and German has 2.3 million articles. (Cebuano and Swedish have a large number of machine generated articles.) In fact, the top nine languages alone hold more than half of all articles across the Wikipedia language editions — and if you take the bottom half of all Wikipedias ranked by size, they together wouldn’t have 10% of the number of articles in the English Wikipedia.

It is not just the sheer number of articles that differ between editions, but their comprehensiveness does as well: the English Wikipedia article on Frankfurt has a length of 184,686 characters, a table of contents spanning 87 sections and subsections, 95 images, tables and graphs, and 92 references — whereas the Hausa Wikipedia article states that it is a city in the German state of Hesse, and lists its population and mayor. Hausa is a language spoken natively by 40 million people and as a second language by another 20 million.

It is not always the case that the large Wikipedia language editions have more content on a topic. Although readers often consider large Wikipedias to be more comprehensive, local Wikipedias may frequently have more content on topics of local interest: the English Wikipedia knows about the Port of Calara?i that it is one of the largest Romanian river ports, located at the Danube near the town of Calara?i — and that’s it. The Romanian Wikipedia on the other hand offers several paragraphs of content about the port.

The topics covered by the different Wikipedias also overlap less than one would initially assume. English Wikipedias has 5.8 million articles, German has 2.2 million articles — but only 1.1 million topics are covered by both Wikipedias. A full 1.1 million topics have an article in German — but not in English. The top ten Wikipedias by activity — each of them with more than a million articles — have articles on only hundred thousand topics in common. 18 million topics are covered by articles in the different language Wikipedias — and English only covers 31% of these.

Besides coverage, there is also the question of how up to date the different language editions are: in June 2018, San Francisco elected London Breed as its new mayor. Nine months later, in March 2019, I conducted an analysis of who the mayor of San Francisco was, according to the different language versions of Wikipedia. Of the 292 language editions, a full 165 had a Wikipedia article on San Francisco. Of these, 86 named the mayor. The good news is that not a single Wikipedia lists a wrong mayor — but the vast majority are out of date. English switched the minute London Breed was sworn in. But 62 Wikipedia language editions list an out-of-date mayor — and not just the previous mayor Ed Lee, who became mayor in 2011, but also often Gavin Newsom (2004-2011), and his predecessor, Willie Brown (1996-2004). The most out-of-date entry is to be found in the Cebuano Wikipedia, who names Dianne Feinstein as the mayor of San Francisco. She had that role after the assassination of Harvey Milk and George Moscone in 1978, and remained in that position for a decade in 1988 — Cebuano was more than thirty years out of date. Only 24 language editions had listed the current mayor, London Breed, out of the 86 who listed the name at all.

<p>The events after the death of Ed Lee until London Breed became mayor on top. On bottom, at what point a given Wikipedia switched.</p>

The events after the death of Ed Lee until London Breed became mayor on top. On bottom, at what point a given Wikipedia switched.

An even more important metric for the success of a Wikipedia are the number of contributors: English has more than 31,000 active contributors — three out of seven active Wikimedians are active on the English Wikipedia. German, the second most active Wikipedia community, already only has 5,500 active contributors. Only eleven language editions have more than a thousand active contributors — and more than half of all Wikipedias have fewer than ten active contributors. To assume that fewer than ten active contributors can write and maintain a comprehensive encyclopedia in their spare time is optimistic at best. These numbers basically doom the mission of the Wikimedia movement to realize a world where everyone can contribute to the sum of all knowledge.

Enter Wikidata

Wikidata was launched in 2012 and offers a free, collaborative, multilingual, secondary database, collecting structured data to provide support for Wikipedia, Wikimedia Commons, the other wikis of the Wikimedia movement, and to anyone in the world. Wikidata contains structured information in the form of simple claims, such as “San Francisco — Mayor — London Breed”, qualifiers, such as “since — July 11, 2018”, and references for these claims, e.g. a link to the official election results as published by the city.

<p>The statement in Wikidata about London Breed being mayor of San Francisco.</p>

The statement in Wikidata about London Breed being mayor of San Francisco.

One of these structured claims would be on the Wikidata page about San Francisco and state the mayor, as discussed earlier. The individual Wikipedias can then query Wikidata for the current mayor. Of the 24 Wikipedias that named the current mayor, eight were current because they were querying Wikidata. I hope to see that number go up. Using Wikidata more extensively can, in the long run, allow for more comprehensive, current, and accessible content while decreasing the maintenance load for contributors.

Wikidata was developed in the spirit of the Wikipedia’s increasing drive to add structure to Wikipedia’s articles. Examples of this include the introduction of infoboxes as early as 2002, a quick tabular overview of facts about the topic of the article, and categories in 2004. Over the year, the structured features became increasingly intricate: infoboxes moved to templates, templates started using more sophisticated MediaWiki functions, and then later demanded the development of even more powerful MediaWiki features. In order to maintain the structured data, bots were created, software agents that could read content from Wikipedia or other sources and then perform automatic updates to other parts of Wikipedia. Before the introduction of Wikidata, bots keeping the language links between the different Wikipedias in sync, easily contributed 50% and more of all edits.

Wikidata allowed for an outlet to many of these activities, and relieved the Wikipedias of having to run bots to keep language links in sync or of massive infobox maintenance tasks. But one lesson I learned from these activities is that I can trust the communities with mastering complex workflows spread out between community members with different capabilities: in fact, a small number of contributors working on intricate template code and developing bots can provide invaluable support to contributors who more focus on maintaining articles and contributors who write large swaths of prose. The community is very heterogeneous, and the different capabilities and backgrounds complement each other in order to create Wikipedia.

However, Wikidata’s structured claims are of a limited expressivity: their subject always must be the topic of the page, every object of a statement must exist as its own item and thus page in Wikidata. If it doesn’t fit in the rigid data model of Wikidata, it simply cannot be captured in Wikidata — and if it cannot be captured in Wikidata, it cannot be made accessible to the Wikipedias.

For example, let’s take a look at the following two sentences from the English Wikipedia article on Ontario, California:

“To impress visitors and potential settlers with the abundance of water in Ontario, a fountain was placed at the Southern Pacific railway station. It was turned on when passenger trains were approaching and frugally turned off again after their departure.”

There is no feasible way to express the content of these two sentences in Wikidata – the simple claim and qualifier structure that Wikidata supports can not capture the subtle situation that is described here.

An Abstract Wikipedia

I suggest that the Wikimedia movement develop an Abstract Wikipedia, a Wikipedia in which the actual textual content is being represented in a language-independent manner. This is an ambitious goal — it requires us to push the current limits of knowledge representation, natural language generation, and collaborative knowledge construction by a significant amount: an Abstract Wikipedia must allow for:

  1. relations that connect more than just two participants with heterogeneous roles.
  2. composition of items on the fly from values and other items.
  3. expressing knowledge about arbitrary subjects, not just the topic of the page.
  4. ordering content, to be able to represent a narrative structure.
  5. expressing redundant information.

Let us explore one of these requirements, the last one: unlike the sentences of a declarative formal knowledge base, human language is usually highly redundant. Formal knowledge bases usually try to avoid redundancy, for good reasons. But in a natural language text, redundancy happens frequently. One example is the following sentence:

“Marie Curie is the only person who received two Nobel Prizes in two different sciences.”

The sentence is redundant given a list of Nobel Prize award winners and their respective disciplines they have been awarded to — a list that basically every large Wikipedia will contain. But the content of the given sentence nevertheless appears in many of the different language articles on Marie Curie, and usually right in the first paragraph. So there is obviously something very interesting in this sentence, even though the knowledge expressed in this sentence is already fully contained in most of the Wikipedias it appears in. This form of redundancy is common place in natural language — but is usually avoided in formal knowledge bases.

The technical details of the Abstract Wikipedia proposal are presented in (Vrandecic, 2018). But the technical architecture is only half of the story. Much more important is the question whether the communities can meet the challenges of this project?

Wikipedia and Wikidata have shown that the communities are capable to meet difficult challenges: be it templates in Wikipedia, or constraints in Wikidata, the communities have shown that they can drive comprehensive policy and workflow changes as well as the necessary technological feature development. Not everyone needs to understand the whole stack in order to make a feature such as templates a crucial part of Wikipedia.

The Abstract Wikipedia is an ambitious future project. I believe that this is the only way for the Wikimedia movement to achieve its goal, short of developing an AI that will make the writing of a comprehensive encyclopedia obsolete anyway.

A plea for knowledge diversity?

When presenting the idea of the Abstract Wikipedia, the first question is usually: will this not massively reduce the knowledge diversity of Wikipedia? By unifying the content between the different language editions, does this not force a single point of view on all languages? Is the Abstract Wikipedia taking away the ability of minority language speakers to maintain their own encyclopedias, to have a space where, for example, indigenous speakers can foster and grow their own point of view, without being forced to unify under the western US-dominated perspective?

I am sympathetic with the intent of this question. The goal of this question is to ensure that a rich diversity in knowledge is retained, and to make sure that minority groups have spaces in which they can express themselves and keep their knowledge alive. These are, in my opinion, valuable goals.

The assumption that an Abstract Wikipedia, from which any of the individual language Wikipedias can draw content from, will necessarily reduce this diversity, is false. In fact, I believe that access to more knowledge and to more perspectives is crucial to achieve an effective knowledge diversity, and that the currently perceived knowledge diversity in different language projects is ineffective at best, and harmful at worst. In the rest of this essay I will argue why this is the case.

Language does not align with culture

First, it is wrong to use language as the dimension along which to draw the demarcation line between different content if the Wikimedia movement truly believes that different groups should be able to grow and maintain their own encyclopedias.

In case the Wikimedia movement truly believes that different groups or cultures should have their own Wikipedias, why is there only a single Wikipedia language edition for the English speakers from India, England, Scotland, Australia, the United States, and South Africa? Why is there only one Wikipedia for Brazil and Portugal, leading to much strife? Why are there no two Wikipedias for US Democrats and Republicans?

The conclusion is that the Wikimedia movement does not believe that language is the right dimension to split knowledge — it is a historical decision, driven by convenience. The core Wikipedia policies, vision, and mission are all geared towards enabling access to the sum of all knowledge to every single reader, no matter what their language, and not toward capturing all knowledge and then subdividing it for consumption based on the languages the reader is comfortable in.

The split along languages leads to the problem that it is much easier for a small language community to go “off the rails” — to either, as a whole, become heavily biased, or to adopt rules and processes which are problematic. The fact that the larger communities have different rules, processes, and outcomes can be beneficial for Wikipedia as a whole, since they can experiment with different rules and approaches. But this does not seem to hold true when the communities drop under a certain size and activity level, when there are not enough eyeballs to avoid the development of bad outcomes and traditions. For one example, the article about skirts in the Bavarian Wikipedia features three upskirt pictures, one porn actress, an anime screenshot, and a video showing a drawing of a woman with a skirt getting continuously shorter. The article became like this within a day or two of its creation, and, even though it has been edited by a dozen different accounts, has remained like this over the last seven years. (This describes the state of the article in April 2019 — I hope that with the publication of this essay, the article will finally be cleaned up).

A look on some south Slavic language Wikipedias

Second, a natural experiment is going on, where contributors that are more separated by politics than language differences have separate Wikipedias: there exist individual Wikipedia language editions for Croatian, Serbian, Bosnian, and Serbocroatian. Linguistically, the differences between the dialects of Croatian are often larger than the differences between standard Croatian and standard Serbian. Particularly the existence of the Serbocroatian Wikipedia poses interesting questions about these delineations.

Particularly the Croatian Wikipedia has turned to a point of view that has been described as problematic. Certain events and Croat actors during the 1990s independence wars or the 1940s fascist puppet state might be represented more favorably than in most other Wikipedias.

Here are two observations based on my work on south Slavic language Wikipedias:

First, claiming that a more fascist-friendly point of view within a Wikipedia increases the knowledge diversity across all Wikipedias might be technically true, but is practically insufficient. Being able to benefit from this diversity requires the reader to not only be comfortable reading several different languages, but also to engage deeply enough and spend the time and interest to actually read the article in different languages, which is mostly a profoundly boring exercise, since a lot of the content will be overlapping. Finding the juicy differences is anything but easy, especially considering that most readers are reading Wikipedia from mobile devices, and are just looking to satisfy a quick information need from a source whose curation they trust.

Most readers will only read a single language version of an article, and thus any diversity that exists across different language editions is practically lost. The sheer existence of this diversity might even be counterproductive, as one may argue that the communities should not spend resources on reflecting the true diversity of a topic within each individual language. This would cement the practical uselessness of the knowledge diversity across languages.

Second, many of the same contributors that write the articles with a certain point of view in the Croatian Wikipedia, also contribute on the English Wikipedia on the articles about the same topics — but there they suddenly are forced and able to compromise and incorporate a much wider variety of points of view. One might hope the contributors would take the more diverse points of view and migrate them back to their home Wikipedias — but that is often not the case. If contributors harbor a certain point of view (and who doesn’t?) it often leads to a situation where they push that point of view as much as they can get away with in each of the projects.

It has to be noted that the most blatant digressions from a neutral point of view in Wikipedias like the Croatian Wikipedia will not be found in the most central articles, but in the large periphery of articles surrounding these central articles which are much harder to keep an eye on.

Abstract Wikipedia and Knowledge diversity

The Abstract Wikipedia proposal does not require any of the individual language editions to use it. Each language community can decide for each article whether to fall back on the Abstract Wikipedia or whether to create their own article in their language. And even that decision can be more fine grained: a contributor can decide for an individual article to incorporate sections or paragraphs from the Abstract Wikipedia.

This allows the individual Wikipedia communities the luxury to entirely concentrate on the differences that are relevant to them. I distinctly remember that when I started the Croatian Wikipedia: it felt like I had the burden to first write an article about every country in the world before I could write the articles I cared about, such as my mother’s home village — because how could anyone defend a general purpose encyclopedia that might not even have an article on Nigeria, a country with a population of a hundred million, but one on Donji Humac, a village with a population of 157? Wouldn’t you first need an article on all of the chemical elements that make up the world before you can write about a local food?

The Abstract Wikipedia frees a language edition from this burden, and allows each community to entirely focus on the parts they care about most — and to simply import the articles from the common source for the topics that are less in their focus. It allows the community to make these decisions. As the communities grow and shift, they can revisit these decisions at any time and adapt them.

At the same time, the Abstract Wikipedia makes these differences more visible since they become explicit. Right now there is no easy way to say whether the fact that Dianne Feinstein is listed as the Mayor of San Francisco in the Cebuano Wikipedia is due to cultural particularities of the Cebuano language communities or not. Are the different population numbers of Frankfurt in the different language editions intentional expressions of knowledge diversity? With an Abstract Wikipedia, the individual communities could explicitly choose which articles to create and maintain on their own, and at the same time remove a lot of unintentional differences.

By making these decisions more explicit, it becomes possible to imagine an effective workflow that observes these intentional differences, and sets up a path to integrate them into the common article in the Abstract Wikipedia. Right now, there are 166 different language versions of the article on the chemical element Helium — it is basically impossible for a single person to go through all of them and find the content that is intentionally different between them. With an Abstract Wikipedia, which contains the common shared knowledge, contributors, researchers, and readers can actually take a look at those articles that intentionally have content that replaces or adds to the commonly shared one, assess these differences, and see if contributors should integrate the differences in the shared article.

The differences in content may be reflecting difference in policies, particularly in policies of notability and reliability. Whereas on first glance it might seem that the Abstract Wikipedia might require unified notability and reliability requirements across all Wikipedias, this is not the case: due to the fact that local Wikipedias can overlay and suppress content from the Abstract Wikipedias, they can adjust their Wikipedias based on their own rules. And the increased visibility of such decisions will lead to easier identify biases, and hopefully also to updated rules to reduce said bias.

A new incentive infrastructure

The Abstract Wikipedia will evolve the incentive infrastructure of Wikipedia.

Presently, many underrepresented languages are spoken in areas that are multilingual. Often another language spoken in this area is regarded as a high-prestige language, and is thus the language of education and literature, whereas the underrepresented language is a low-prestige language. So even though the low-prestige language might have more speakers, the most likely recruits for the Wikipedia communities, people with education who can afford internet access and have enough free time, will be able to contribute in both languages.

In which language should I contribute? If I write the article about my mother’s home town in Croatian, I make it accessible to a few million people. If I write the article about my mother’s home town in English, it becomes accessible to more than a hundred times as many people! The work might be the same, but the perceived benefit is orders of magnitude higher: the question becomes, do I teach the world about a local tradition, or do I tell my own people about their tradition? The world is bigger, and thus more likely to react, creating a positive feedback loop.

This cannibalizes the communities for local languages by diverting them to the English Wikipedia, which is perceived as the global knowledge community (or to other high-prestige languages, such as Russian or French). This is also reflected in a lot of articles in the press and in academic works about Wikipedia, where the English Wikipedia is being understood as the Wikipedia. Whereas it is known that Wikipedia exists in many other languages, journalists and researchers are, often unintentionally, regarding the English Wikipedia as the One True Wikipedia.

Another strong impediment to recruiting contributors to smaller Wikipedia communities is rarely explicitly called out: it is pretty clear that, given the current architecture, these Wikipedias are doomed in achieving their mission. As discussed above, more than half of all Wikipedia language editions have fewer than ten active contributors — and writing a comprehensive, up-to-date Wikipedia is not an achievable goal with so few people writing in their free time. The translation tools offered by the Wikimedia Foundation can considerably help within certain circumstances — but for most of the Wikipedia languages, automatic translation models don’t exist and thus cannot help the languages which would need it the most.

With the Abstract Wikipedia though, the goal of providing a comprehensive and current encyclopedia in almost any language becomes much more tangible: instead of taking on the task of creating and maintaining the entire content, only the grammatical and lexical knowledge of a given language needs to be created. This is a far smaller task. Furthermore, this grammatical and lexical knowledge is comparably static — it does not change as much as the encyclopedic content of Wikipedia, thus turning a task that is huge and ongoing into one where the content will grow and be maintained without the need of too much maintenance by the individual language communities.

Yes, the Abstract Wikipedia will require more and different capabilities from a community that has yet to be found, and the challenges will be both novel and big. But the communities of the many Wikimedia projects have repeatedly shown that they can meet complex challenges with ingenious combinations of processes and technological advancements. Wikipedia and Wikidata have both demonstrated the ability to draw on technologically rather simple canvasses, and create extraordinary rich and complex masterpieces, which stand the test of time. The Abstract Wikipedia aims to challenge the communities once again, and the promise this time is nothing else but to finally be able to reap the ultimate goal: to allow every one, no matter what their native language is, to share in the sum of all knowledge.

Acknowledgements

Thanks to the valuable suggestions on improving the article to Jamie Taylor, Daniel Russell, Joseph Reagle, Stephen LaPorte, and Jake Orlowitz.

Header Image: Created by Bleeptrack, https://commons.wikimedia.org/wiki/File:Large_Wikidata_Pattern.png

Bibliography

  • Bao, Patti, Brent J. Hecht, Samuel Carton, Mahmood Quaderi, Michael S. Horn and Darren Gergle. “Omnipedia: bridging the wikipedia language gap.” in Proceedings of the Conference on Human Factors in Computing Systems (CHI 2012), edited by Joseph A. Konstan, Ed H. Chi, and Kristina Höök. Austin: Association for Computing Machinery, 2012: 1075-1084.
  • Eco, Umberto. The Search for the Perfect Language (the Making of Europe). La ricerca della lingua perfetta nella cultura europea. Translated by James Fentress. Oxford: Blackwell, 1995 (1993).
  • Graham, Mark. “The Problem With Wikidata.” The Atlantic, April 6, 2012. https://www.theatlantic.com/technology/archive/2012/04/the-problem-with-wikidata/255564/
  • Hoffmann, Thomas and Graeme Trousdale, “Construction Grammar: Introduction”. In The Oxford Handbook of Construction Grammar, edited by Thomas Hoffmann and Graeme Trousdale, 1-14. Oxford: Oxford University Press, 2013.
  • Kaffee, Lucie-Aimée, Hady ElSahar, Pavlos Vougiouklis, Christophe Gravier, Frédérique Laforest, Jonathon S. Hare and Elena Simperl. “Mind the (Language) Gap: Generation of Multilingual Wikipedia Summaries from Wikidata for Article Placeholders.” in Proceedings of the 15th European Semantic Web Conference (ESWC 2018), edited by Aldo Gangemi, Roberto Navigli, Marie-Esther Vidal, Pascal Hitzler, Raphaël Troncy, Laura Hollink, Anna Tordai, and Mehwish Alam. Heraklion: Springer, 2018: 319-334.
  • Kaffee, Lucie-Aimée, Hady ElSahar, Pavlos Vougiouklis, Christophe Gravier, Frédérique Laforest, Jonathon S. Hare and Elena Simperl. “Learning to Generate Wikipedia Summaries for Underserved Languages from Wikidata.” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2, edited by Marilyn Walker, Heng Ji, and Amanda Stent. New Orleans: ACL Anthology, 2018: 640-645.
  • Schindler, Mathias and Denny Vrandecic. “Introducing new features to Wikipedia: Case studies for Web Science.” IEEE Intelligent Systems 26, no. 1 (January-February 2011): 56-61.
  • Vrandecic, Denny. “Restricting the World.” Wikimedia Deutschland Blog. February 22, 2013. https://blog.wikimedia.de/2013/02/22/restricting-the-world/
  • Vrandecic, Denny and Markus Krötzsch. “Wikidata: A Free Collaborative Knowledgebase.” Communications of the ACM 57, no. 10 (October 2014): 78-85. DOI 10.1145/2629489.
  • Kaljurand, Kaarel and Tobias Kuhn. “A Multilingual Semantic Wiki Based on Attempto Controlled English and Grammatical Framework.” in Proceedings of the 10th European Semantic Web Conference (ESWC 2013), edited by Philipp Cimiano, Oscar Corcho, Valentina Presutti, Laura Hollink, and Sebastian Rudolph. Montpellier: Springer, 2013: 427-441.
  • Milekic, Sven. “Croatian-language Wikipedia: when the extreme right rewrites history.” Osservatorio Balcani e Caucaso, September 27, 2018. https://www.balcanicaucaso.org/eng/Areas/Croatia/Croatian-language-Wikipedia-when-the-extreme-right-rewrites-history-190081
  • Ranta, Aarne. Grammatical Framework: Programming with Multilingual Grammars. Stanford: CSLI Publications, 2011.
  • Vrandecic, Denny. “Towards a multilingual Wikipedia,” in Proceedings of the 31st International Workshop on Description Logics (DL 2018), edited by Magdalena Ortiz and Thomas Schneider. Phoenix: Ceur-WS, 2018.
  • Wierzbicka, Anna. Semantics: Primes and Universals. Oxford: Oxford University Press, 1996.
  • Wikidata Community: “Lexicographical data.” Accessed June 1, 2019. https://www.wikidata.org/wiki/Wikidata:Lexicographical_data
  • Wulczyn, Ellery, Robert West, Leila Zia and Jure Leskovec. “Growing Wikipedia Across Languages via Recommendation.” in Proceedings of the 25th International World-Wide Web Conference (WWW 2016), edited by Jaqueline Bourdeau, Jim Hendler, Roger Nkambou, Ian Horrocks, and Ben Y. Zhao. Montréal: IW3C2, 2016: 975-985.

Kopfzerbrechen Nr. 3: Genealogien

Genealogien durchdringen alle Arbeitsgebiete. Wer Bibliothekskataloge durchsucht, erhält Verlagsangaben und bleibt im Dunkeln darüber, wie diese Angaben zusammenhängen.

Genealogien der Geschäftsführung

Dem Buchhandelshistoriker ist klar, dass er hier genealogisch denken muss: Die städtischen Behörden vergaben Lizenzen für den Druck, Verlag und/oder Verkauf von Büchern. Die Lizenzen wurden innerfamiliär weitergegeben vom Vater auf den Sohn oder, bei ausbleibenden Erben, auf die Witwe und die Kinder, bis jemand in das Geschäft einheiratete oder es von auswärts kommend erwarb.

Neue Lizenzen wurden selten eröffnet. Das Interesse des Hofes konnte hier innserstädtische Interessen an der Vermeidung von Konkurrenz außer Kraft setzen. Eine Geschäftsteilung konnte eine Differenzierung in verschiedene Lizenzen mit sich bringen – fortan bestand ein Unternehmen als reine Druckerei, und eines als reines Verlagsgeschäft mit Buchhandlung.

Dutzende von Namen über drei Jahrhunderte ordnen sich, sobald man die Genealogie der weitergegebenen Rechte erfasst, zu wenigen Strängen pro Stadt – das kann wie folgt für München oder Gotha aussehen:

Vergleichbare Darstellungen der Geschäftsbeziehungen sind bereits Teil des ungeborgenen Katalogwissens. Die Kataloge könnten theoretisch auf die Nahtstellen hinweisen, an denen vermutlich ein Wechsel stattfand, ein Name ausläuft, ein anderer anhebt. Letztlich erfordert die Rekonstruktion der Zusammenhänge jedoch eigenes Wissen über Familienbeziehungen und die Zusammenhänge der städtischen Gewerbeakten.

Auf der Ebene der Software benötigte man die passenden Darstellungsoptionen. Vor allem aber braucht man Schnittstellen, an denen Benutzer die Verbindungen herstellen und die Genealogie ordnen können.

Genetische Verhältnisse: Stemmata

Viel komplizierter liegen die Verhältnisse bei den genetischen Zusammenhängen zwischen Buchausgaben. Manche Kataloge scheitern hier schon im Ansatz wie der hinter Google Books. Man versuche etwa, die Nummern eines Journals aus dem 18. Jahrhundert, dessen Digitalisate irgendwie im System stecken, sich der Reihenfolge nach anzeigen zu lassen. Google hilft einem hier mit einem vagen Angebot weiter nach dem Motto „Leser, die dies lasen, könnten auch diese Links interessant finden“.

Doch auch ausgefeilte Kataloge wie der ESTC, das Verzeichnis aller englischen Titel der Jahre 1473 bis 1800, weisen hier rasch unüberschaubar werdende Gebiete auf – dann etwa, wenn minimale Varianten von einer Ausgabe, verschiedene Auflagen und zudem einzelne Teilbände auf ein Jahr fallen wie im Fall der Katalogangabe für Delarivier Manleys Atalantis:

ESTC Suche: “Atalantis”. Auch nach der chronologischen Sortierung bleibt unklar, hinter welchem Link der erste Text liegt.

Man wünschte sich hier nicht minder eine Option der genealogischen Strukturierung: Der erste Band erschien im Mai 1709; im Juli musste er nachgedruckt werden. Band zwei folgte im Oktober. 1710 brachte die Autorin ein Werk heraus, das erst einmal einen eigenen Titel trug: die Memoirs of Europe. Ein halbes Jahr später folgte davon Band zwei. Ab 1715/16 wurden diese Bände in Neuausgaben der Atalantis als deren Folgen III und IV notiert.

Dem Markterfolg wurden ab 1711 auch noch ganz andere Bücher untergeschoben: 1705 war erstmals eine Queen Zarah mit ähnlicher Skandalgeschichte erschienen. Deren zweite Ausgabe kommt 1711 „by way of appendix to the New Atalantis“ heraus. Ein weiterer älterer Titel macht denselben Sprung in den Komplex und vier neue wollen unverzüglich in ihn hinein. Siehe hier die geschlossenere genetische Darstellung: http://pierre-marteau.com/library/e-1709-0004.html.

Der ESTC ist im Kern ein Nationalkatalog. Das wird deutlicher, sobald man die weiteren Titel auf dem europäischen Markt nachweisen will – dem versagt sich der ESTC. In den Niederlanden erfolgte 1713 die Übersetzung der ersten zwei Bände ins Französische. Eine nachweisbare, kondensierte deutsche Fassung datiert vermutlich von 1714 und übersetzt nach der französischen Ausgabe.

Für die beschriebenen genetischen Beziehungen bräuchte man Stemmata, wie sie in der Handschriftenkunde verbreitet sind. Das Komplizierte ist hier, dass die Beziehungen zwischen Titeln immer über das ganze Schema hinweg verlaufen können. Ein Verlag mag bei einer Neuausgabe nach der letzten ihm greifbare Ausgabe setzen. Er könnte jedoch auch auf die Erstausgabe zurückgreifen, oder eine von der Autorin korrigierte spätere. Die moderne „kritische“ Ausgabe wird sich für eine Leitausgabe entscheiden, und Varianten (wenn vorhanden) mit dem Manuskript der Autorin und den ersten und letzten Ausgaben, die sie selbst beeinflusste, erfassen.

Stemma für die Manuskripte des ‘Pseudo-Apuleius Herbarius’ aus Ernst Howald and Henry E. Sigerist. Antonii Musa De herba vettonica (Leipzig 1927). https://commons.wikimedia.org/wiki/File:Howald-sigerist.png

Wieder wäre einem mit einer Schnittstelle geholfen, die es erlaubte, verschiedene genetische Beziehungen zwischen Ausgaben herzustellen, die dann in einer Visualisierung zum Zuge kämen.

Entwicklungen

Natürlich hofft man auf die Software, die selbst rekonstruiert, wie sich die Dinge ordnen – darum geht es im Umgang mit „Big Data“. Das Spannende an der Wikidata-Software ist jedoch im selben Moment, dass der Benutzer mit ihr sein Wissen viel schneller und härter einbringen und damit die großen Schneisen der Diskussion gezielt und diskutierbar schlagen könnte.



oder:

..von hier: https://www.edwardtufte.com/bboard/q-and-a-fetch-msg?msg_id=0000yO

Siehe auch

Faktenbelege, die “Original Research” zulassen

See https://factgrid-tools.geschichte.uni-halle.de/blog/archives/978 for a revised version in English.

Bei der Arbeit mit der Wikidata Software werden Anwender aus dem wissenschaftlichen Feld andere Anforderungen an den Beleg von Informationen stellen.

In Wikipedia-Projekten (und das schließt Wikidata ein) ist der Quellenbeleg zwar die grundlegende Anforderung; da man primär “enzyklopädisch relevantes” Wissen sammeln will, genügt hier jedoch der Verweis auf eine Publikation, die das Faktum so notiert. Bei widersprüchlichen Informationen werden die verschiedenen Angaben kommentarlos nebeneinander gestellt. Wikidata sieht in der Folge in der Regel einen einfachen Beleg vor: die bibliographische Angabe oder ein Link.

„Original Research“ ist in Wikipedia ausgeschlossen. Der Grund dafür ist strukturell: Wikipedia-Autoren bleiben zum guten Teil anonym – das erforderte die Beschränkung ihrer Beiträge auf eine Weitergabe öffentlich verfügbaren, von externen Autoritäten zitierbaren Wissens.

Im Moment, in dem Forschung dieselbe Software benutzen will, muss der Beiträger sich mit seinem Namen für den Befund verbürgen können. Der Quellenbeleg ist dabei nicht mehr unbedingt eine bereits verfügbare Publikation. Genauso gut kann eine Archivalie oder eine Schlussfolgerung zum Beleg werden. Die Datenbank muss nun Fakten als Forschungsbeiträge präsentieren können – ein beliebiges Faktum muss als eine Mikropublikation zitierbar sein. Der Forscher, der sich für den Befund verbürgt, muss namentlich genannt werden, seine Erwägungen müssen im komplexen Fall einer Schlussfolgerung mit dem Befund verfügbar werden. Das Einstelldatum des Befunds macht aus dem Beitrag eine wissenschaftliche Publikation. Forscher können andernfalls nicht riskieren, Befunde vor einer anderweitigen Publikation hier bereits allgemein verfügbar zu machen.

Belege im FactGrid werden zu diesem Zweck sich wohl am besten in einem Modul mit Pulldown-Menüs und Eingabefeldern verwalten lassen. Vielleicht wird man über eine Option nachdenken, wie man eine solche Fußnote arbeitssparend duplizieren kann.

1. Ausweis der Forschungsleistung

  1. Die Information liegt öffentlich vor
    • Eingabefeld für Publikation und Seitenangabe, respektive Datenbanknennung und Link
    • Anweisung, eine Kurzpublikation mit DOI zu erstellen [Daniel Mietchens Tipp]
  2. Die Information wird hier erstmals veröffentlicht
    • Eingabefeld für die zu zitierenden Forscher mit Einstelldatum (Default: automatische Signatur)

2. Quelle

  1. Die Quelle ist öffentlich zugänglich
    • Eingabefeld für X-Nummer in dieser Datenbank mit dortigem Archivnachweis und wenn möglich mit Scan
  2. Die Quelle ist öffentlich unzugänglich
    • Eingabefeld für X-Nummer in dieser Datenbank mit dortigen weiteren Informationen und eventuellem Scan
  3. • persönliches Wissen des Veröffentlichenden

3. Status der Information

  1. solide, laut Dokument
  2. begründete Schlussfolgerung
  3. Arbeitshypothese
  4. vage vorliegende Information
  5. widerlegt
  6. privates Wissen/ Familienwissen
  7. allgemeine Annahme
  8. Behauptung in einem fiktionalen Rahmen
  9. Behauptung in religiöser Annahme

4. Diskussionsfeld

5. Zitierlink, das die Informationen zu einer Fußnote zusammenfasst

Es geht mit diesen Optionen darum, den noch nirgends publizierten Befund notierbar zu machen und an eine Diskussion und Bewertung zu koppeln. Man kann im selben Moment zulassen, dass Beiträger selbst privates Wissen (etwa um persönliche genealogische Beziehungen) einbringen, denn es bleibt als solches in seiner Angreifbarkeit wie seinen ganz eigenen Qualität ausgewiesen.

Auf dem Weg zur automatisch generierten Fußnote in Wikipedia

Aus den Angaben sollten sich letztlich automatisiert Fußnoten in beliebigen Sprachen generieren lassen: Goethe wurde am 28. August 1749 geboren – die Fußnote dazu notiert, aus welchem Dokument oder welcher späteren Quelle wir das wissen und führt im brisanten Fall zu einer Diskussionsseite, auf der Forscher sich über die Angabe einigten.

Ausgangsüberlegungen

1| Ausgangslage

Für die Arbeit an der Gotha Illuminati Research Base benutzten wir in den letzten Jahren mit Erfolg ein ganz konventionelles Wiki. Das war besonders beim Materialmix von Vorteil, den wir zu verwalten haben:

Das Wiki bewährte sich, da es ohne Vorplanung entlang unserer Befunde wachsen konnte. Unsere Forschungsarbeit griff dank der praktischen Ressource auf die breite Materialbasis aus.

2| Unser Projekt geriet an technische Grenzen —

Metadaten zu einer typischen Archivalie wie sie gleichlautend parallel in mehreren Tabellen verwaltet werden müssen. Link in die Seite
  • da die Anzahl von Listen in ihm wuchs, die letztlich auf dieselben Fakten zugreifen. Ändert sich ein einzelner Befund, muss sichergestellt werden, dass alle Listen die neue Datenlage zeigen.
  • da uns bei den Biographien besser mit automatisch generierten CVs gedient wäre – mit gerne auch langen tabellarischen Listungen etwa aller Wohnsitze, aller Ausbildungsetappen, aller Arbeitsverhältnisse, aller Korrespondenzen und so fort. Die biographischen Seiten des Wikidata Reasonators sind hier weitgehend mustergültig. (Man würde sich eine klarere Strukturierung wünschen und bei mehr Einträgen eine Option der vollen Anzeige auf Wunsch – siehe die eigenen Gedanken zu Biographien und Eingabeschablonen in diesem Blog.)
  • da wir mit der bisherigen Verwaltung einzelner Seiten nicht in die Lage kommen, Informationen automatisch zu generieren und etwa zu Netzwerk-Analysen und Bewegungsprofilen auf Landkarten zusammenzuführen.
  • da wir gerne andere Projekte in dieselbe Ressource hineinarbeiten lassen würden, was vor allem im Feld der reinen Datenlage interessant wird.
  • da das Interesse an unserer Arbeit vor allem im anglophonen Ausland liegt, den wir aber nur mit einem erheblichen Arbeitsaufwand ansprechen könnten. Hier beeindruckte erneut der Reasonator bei Vorführungen in seiner Fähigkeit, die Sprachen zu wechseln.
Die automatisch generierte Seite zu Ernst II. von Sachsen-Gotha-Altenburg aus dem Wikidata Reasonator
https://tools.wmflabs.org/reasonator/?&q=213698

3| Wie die Wikidata Software einsetzen?

Die Wikidata Software dürfte die perfekte Problemlösung sein, da sie sich in derselben Problemstellung unter einem sehr viel breiteren Interesse entwickelte und verspricht, in Entwicklung zu bleiben.

Es schien uns nach der ersten Sondierung im Januar 2016 nicht ratsam, Wissenschaftler direkt in den von Wikimedia betreuten Wikidata-Pool hineinarbeiten zu lassen. Die Forschungsressource wird als Ort von “original research” manches anders handhaben, und gerade dass sie manches anders handhabt, wäre für Wikidata als Nutznießer einer Kooperation von Vorteil. Die Lösung wäre hier eine eigene, die Wikidata-Software benutzende Ressource – im Folgenden das FactGrid – die mit Wikidata eine Datenpartnerschaft einginge.

Zentrale technische Fragen sind dabei von Seiten des wissenschaftlichen Projektes:

  • Sollen wir unser bisheriges Wiki in dieser Datenbank aufgehen lassen, indem wir dessen ganze Materialvielfalt (siehe oben Punkt 1) in die Verwaltung der Datenbank überführen, oder würden wir das FactGrid besser ähnlich wie die verschiedensprachigen Wikipedia-Projekte nutzen, die gemeinsam aus Wikimedia Commons etwa Mediendateien beziehen – nun um Fakten zu verwalten?
  • Wie würden verschiedene Projekte in dieselbe Ressource hineinarbeiten und im jeweils eigenen Forschungsinteresse von ihr profitieren?
  • Wie gestaltet sich die Befüllung mit Daten in größeren Mengen (etwa aus Excel-Tabellen wie dieser zu den Akten der “Schwedenkiste“)? Wie generiert man Eingabeschablonen, um Fragekataloge abzuarbeiten? (Dies sind etwa die Optionen einer Mitgliedschaft im Illuminatenorden). Wie stimmen sich Projekte bei der Nutzung der gemeinsame Ressource ab?
  • Wie beziehen unabhängig voneinander arbeitende Projekte Daten in komplexeren Anwendungen aus der gemeinsamen Ressource (etwa um Netzwerke oder Bewegungsprofile einzelner Personen zu erstellen)?
  • Wie lassen sich komplexere Authentifizierungen von Forschungsbefunden bewerkstelligen? Diese müssen Aktenquellen zulassen und als Forschungsbefunde zitierbar werden. Wir brauchen dabei zudem Bewertung von Befunden durch die Forscher, um etwa “hypothetische” Aussagen zur weiteren Verifikation aussetzen zu können.
  • Wie würde sich ein kontrollierter Datenfluss aus dem FactGrid zu Wikidata organisieren lassen? (Sicherzustellen ist hierbei, dass Daten wie bei uns authentifiziert und bewertet bleiben.)

4| Eine Wikimedia Kooperation: Organisatorische Aspekte

Ein konstantes Ärgernis von Projekten der Digital Humanities ist bislang die Produktion von Insellösungen. Eine Software wird auf das einzelne Projekt zugeschnitten. Dessen Befunde enden mit der zunehmend antiquierten Plattform außerhalb einer weiteren Benutzung.

Das Ziel sollte die Investition in Tools sein (z.B. zur Visualiserung von Netzwerken oder Bewegungsprofilen auf Landkarten), die danach mit der breit verwendeten Software verfügbar werden. Forschungsgelder sollten in die Fortentwicklung bestehender Lösungen fließen, in neue Features, wo bislang immer neue Projekte dasselbe Rad für ihre Software neu erfinden.

Für eine Kooperation zwischen Projekten der Digital Humanities und Wikimedia kann Gotha derzeit günstige Ausgangsbedingungen anbieten: Wir verfügen mit der Gotha Illuminati Research Base über ein Startprojekt immensen öffentlichen Sex-Appeals, und das nun nicht mehr nach einer Struktur suchen muss. Im Kern liegen die Datensätze vor, mit denen Folgeprojekte an eine Feinerschließung aller Illuminatenakten gehen können.

Gleichzeitig können wir auf die nächsten fünf Jahre eine koordinierte weit größere integrative Arbeit mit dem Projekt „Gotha um 1800: Natur-Wissenschaft-Geschichte“ anbieten. Dieses Projekt wird unterschiedliche Institutionen zusammenbringen und sehr unterschiedliche Materialien in neue Zusammenhänge setzen: Informationen zu Wissenschaften um 1800, ortsbezogene Informationen zu einem deutschen Fürstentum, Informationen zu zahllosen Artefakten, Informationen zu Büchern und Zeitschriften aus dem Untersuchungszeitraum 1750 bis 1850, Informationen zu Personen und Personengeflechten, öffentlichen und geheimen Gesellschaften, ortsansässigen und korrespondierenden.

Mehrere Projekte am Forschungszentrum meldeten ihr Interesse an einer gemeinsamen Datenbanklösung an, falls wir eine solche angingen. Wir sind hier in der Lage, langfristig eine Community aufzubauen, die das nötige Know How intern verwalten kann, das regulär mit einzelnen Projekten nach deren Förderung verloren geht.

Eine Kooperationsvereinbarung mit Wikimedia sollte:

  • den Rahmen einer Betreuung abstecken, die die Wissenschaftler vor allem in der Anlaufphase benötigen, bevor sie sich in der Lage sehen, Problemlösungen intern mit hinzukommenden Projekten zu teilen.
  • erfassen, wie Software-Entwicklungen angestoßen werden und gegebenfalls aus eigenen Mitteln der Forschungsprojekte finanziert werden können.
  • sicherstellen, dass Entwicklungen neuer Wikidata-Features als Tools in das Angebot der freien Software gelangen, so dass sie von dort aus die breitere Nutzung der Software in Wissenschaftskreisen interessant machen.
  • Ideen entwickeln, wie die Community, die Wikidata betreut, mit den Forschern, die das FactGrid-aufbauen, in einen für beide Seiten interessanten Austausch gelangen. Dabei wird es vor allem um den Datenfluss zu Wikidata und um praktische Erfahrungen bei der Benutzung der gemeinsamen Software gehen.

Entscheidend wird es sein, in den Vorüberlegungen die gemeinsamen Interessen zu erfassen. Die Wikidata Software ist eine der interessantesten Entwicklungen der letzten Jahre. Sie löst auf spannende Weise Probleme, die viele wissenschaftliche Projekte derzeit mit Datenbank-Softwareangeboten haben. Sie ist auf vielfältige Objekte zugeschnitten, kann große Datenvolumina verwalten, notiert Veränderungen und Bearbeitungen nachvollziehbar und ihre Ressource agiert erfolgreich in allen möglichen Sprachen der Welt – doch ist sie gleichzeitig derzeit hermetisch abgeschlossen.

Die wissenschaftliche Arbeit mit dieser Software verspricht einen Erfahrungsaustausch, von dem die Software profitieren kann. Das Wikidata-Projekt hat dabei langfristig das Potential, Limitationen des Wikipedia-Projektes negieren zu können, jene Limitationen, die Forschenden die Arbeit in Wikipedia schwer machen. Der Schritt in die wissenschaftliche Ressource wird indes auch von Software-Anpassungen abhängen, die gemeinsam mit Wissenschaftlern gefunden und ausgetestet werden müssen. Dies vor allem sollte es auf Wikimedia-Seite interessant machen, Pilotprojekte mit der in den letzten Jahren aufgebauten Software betreut arbeiten zu lassen.

Der Aufbau einer Ressource, die “original research” betreiben und autorisieren kann und der Datenfluss von der wissenschaftlichen Ressource zu Wikidata sollten das Interessenspektrum erweitern.

Links etc.

  • The Gotha Illuminati Research Base. Link
  • Hier im Blog: Die Schwedenkiste digital: die ausstehende Durchleuchtung des Illuminatenordens
  • Metadaten zu den knapp über 6000 Dokumenten der “Schwedenkiste”, des zentralen Aktenbestands im Nachlass des Illuminatenordens. Link
  • Bislang ungenützt: der FactGrid Prototyp mit der Wikidata Software. Noch weiß keiner unserer Wissenschaftler, wo er Daten einfüllen soll und wie er sie dann wieder zu Anwendungen herauszieht.