FactGrid GYIK – Miért használjam a FactGridet a kutatási projektemhez?

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).

in English
auf Deutsch
en français

  1. Mi a FactGrid?
  2. Miért használjam a FactGridet a saját kutatásomhoz?
  3. Miért ne egyből a Wikidatát használjam?
  4. A FactGrid ingyenes – hogy működik ez?
  5. Mihez kezdhetek az unortodox kutatási témákkal?
  6. Milyen segédeszközöket biztosít a szoftver?
  7. Mit tegyek, ha a saját platformomon szeretném megjeleníteni az adatvizualizációm?
  8. A FactGrid CC0-licenc alatt teszi közzé az adatokat – ez azt jelenti, hogy lemondok a kutatásom jogairól?
  9. Mi történik, ha szeretném az adataimmal egy másik platformon folytatni a munkát?
  10. Mi történik, amikor FactGrid-felhasználók a “helyes” dátumról vitatkoznak?
  11. Miért kockáztassam meg az átláthatóságot rögtön a projektem kezdetétől?
  12. Mi kell ahhoz, hogy a FactGrid befogadja a projektem?

Mi a FactGrid?

A FactGrid egy Wikibase-alapú platform történeti adatokkal dolgozó projekteknek számára, amely egyszerre hagyományos wiki és adatbázis. Az oldalon állításokat rögzíthetsz az általad feltöltött vagy téged érdeklő elemekről, majd ezeket szinte bármilyen nyelven tudod használni és megjeleníteni.

A platform szervezője a Gotha Kutatóközpont, a szervert pedig ThULB Jena biztosítja.

Együttműködésben a Wikimédia Németországgal és a Német Nemzeti Könyvtár GND-adatbázisával szeretnénk elhelyezni a platformot mint kutatási adatokra építkező erőforrást a kialakulóban lévő, összekapcsolt Wikibase-oldalak rendszerében.

Miért használjam a FactGridet a saját kutatásomhoz?

A fő érv a FactGrid mellett a verhetetlenül rugalmas szoftver, a Wikibase, amelyet a Wikimédia Németország segítségével, elsődleges felhasználási helyén, a Wikidatán kívül, egy kísérleti projekt keretében implementáltunk:

  • Egy olyan szoftvert keresel, amely gyakorlatilag bármilyen nyelven tud beszélni? Egy platformot, ahol felvihetsz adatokat a saját nyelveden, mások pedig a saját anyanyelvükön olvashatják ugyanezt, és fordítva? Ez a szoftver a Wikibase.
  • Egy olyan szoftverre van szükséged, amivel átlátható módon koordinálhatsz egy egész kutatói csapatot? A Wikibase-zel ez ugyanolyan könnyű, mint a Wikipédia szoftverével, a MediaWikivel.
  • Egy olyan adatbázisszoftvert keresel, amely tud mindent, amire egy digitális bölcsészeti adatbázisnak szüksége lehet: kapcsolatháló-elemzés, térképes megjelenítés, komplex összekapcsolt keresések, megjelenítés többféle idővonalon? Egy szoftver, amely szinte emberi nyelvként működik, és még teljes körű adatbázis szolgáltatással is rendelkezik? A Wikibase ez a szoftver.
  • Szeretnél egy előző projektedből származó adatgyűjteményre építeni? A Wikibase-en lehetséges a nagy mennyiségű, automatizált adatbevitel.
  • Szeretnél biztosra menni, hogy más projektek is hozzáférnek az adataidhoz, és ténylegesen fel is tudják használni azokat? A platformról könnyen letöltheted az összes adatot, hogy offline, Excelben vagy bármilyen más online projektben dolgozhass velük.
  • Szeretnél teljesen új kérdéseket feltenni a kutatásodban? A Wikibase-en bármelyik elemet összekapcsolhatod bármiféle állítással.
  • Aggódsz, hogy mi történik majd az adataiddal miután véget ér a kutatásod finanszírozása? Támaszkodj egy platformra, ahol nem egyedül dolgozol, ami olyan licenc alatt működik, amely lehetővé teszi másoknak is, hogy folytassák a munkát az adataiddal és eszközeiddel.

Ha hosszú távú perspektívát keresel, akkor ezt szeretnénk nyújtani a Német Nemzeti Könyvtárral való együttműködésünkkel. A platform egyik támpillére a GND-adatgyűjtemény lesz, ami által széles körben használható eszközként működhetünk. Továbbá célunk ezzel, hogy fontos szereplőjévé váljunk az összekapcsolt Wikibase-rendszerek kialakuló világának.

Miért ne egyből a Wikidatát használjam?

Ez egy teljesen jogos kérdés. Vannak olyan projektek (amelyek elsősorban csak felhasználják adatokat), amelyekhez a Wikidata megfelelőbb platformot nyújt. Az FH Potsdam “Archivführer zur deutschen Kolonialzeit” nevű projektje remekül illusztrálta annak szépségét, amikor közvetlenül Wikidatára dolgozunk – erről beszélgettünk Uwe Junggal, aki bemutatta, milyen technikai megoldásokat használtak Potsdamban.

Ugyanakkor alapvetően két dolog van, amiket nem fogsz tudni sem a Wikidatán, sem egy GND-hez hasonló platformon csinálni: a Wikimédia-projektek (és a GND) szigorú szabályokkal rendelkeznek arról, hogy nem közölhető saját kutatómunka, és döntéseiket nevezetességi kritériumok alapján hozzák meg, ami nem enged teret tetszőleges adatbázis-elemek létrehozásának vagy tárgyak közötti kísérleti kapcsolatok tesztelésének.

A Wikidata és a GND olyan információkra koncentrálnak, amelyeket már korábban publikáltak és a kutatást nem végző alkalmazottak már közzétett kutatásokból viszik fel az adatokat. Ezeken a platformokon nem tudsz létrehozni munkahipotézisként szolgáló állításokat a kutatásodhoz. Nem hozhatsz létre elemeket kizárólag azzal a céllal, hogy majd statisztikai elemzést végezhess rajtuk a munka egy jóval későbbi szakaszában.

A FactGriden bátorítjuk a platform használatát heurisztikus kutatási eszközként.

  • Létrehozhatsz elemeket az adatbázisban függetlenül attól, milyen relevanciájuk lenne egy enciklopédiában vagy könyvtári katalógusban.
  • Megkockáztathatsz ideiglenes kronológiákat, egyéni feltevéseket kiinduló hipotézisként.
  • Használd a FactGridet nem konvencionális állításokhoz, amelyek jelenleg csak a saját kutatási projekted számára érdekesek – a szoftver lehetővé teszi ezt a fajta szabadságot.
  • Hozz létre adatbázis elemeket, amelyek részletezik, a kutatásod során milyen adatgyűjteményeket módosítottál jelentős mértékben. Ezáltal könnyen benyújthatod ezt az adott elemet mint a kutatásodat összegző “mappát” a téged finanszírozó intézménynek.
  • A platformon megkockáztathatsz bármilyen új tézist, és egy saját adatbázis elemben összegezheted mint “mikro-publikációt”, ezáltal is láthatóvá téve a hozzájárulásod.

A FactGrid ingyenes – hogy működik ez?

A szoftver ingyenesen használható, és folyamatosan fejlesztik a Wikimédia projektek közösségei, illetve a Wikibase-t használó intézmények.

A FactGrid platformot a Gotha Kutatóközpont szolgáltatja az Erfurti Egyetem virtuális szerverén. A német URL évi 36 eurós költséget jelent, ezt a Gotha Kutatóközpont fedezi.

Az összes Wikidata-segédeszköz a felhasználóink rendelkezésére áll. Ezek biztosítják az átlag digitális bölcsészeti projekthez szükséges összes funkciót.

Mivel mind a szoftver, mind az eszközök nyílt forráskóddal rendelkeznek, bármilyen általad kedvelt szoftverrel módosíthatod őket, ha új alkalmazási módra van szükséged.

Ha saját eszközeiddel is hozzájárulsz a nyílt rendszerhez, biztosíthatod, hogy jövőbeli projektek is használhatják és fejleszthetik ezeket.

Amennyiben olyan technikai megoldásokra törekszel, amelyeket később anyagi haszonért értékesíthetsz, a szoftver licence ebben sem fog meggátolni. Szabadon kereskedelmi alapokra helyezhetsz bármit, amit nyílt forráskóddal építettél.

Mihez kezdhetek az unortodox kutatási témákkal?

A Wikidata úttörő adatmodellel rendelkezik. A felhasználó gyakorlatilag csak kapcsolatokat hoz létre Q-számok között (vagy kapcsolatokat Q-számok és időpontok, Q-számok és földrajzi koordináták, Q-számok és médiafájlok, Q-számok és URL-ek között).

A szoftver maga nem tudja, milyen típusú kapcsolatokat hozol létre – ezek szintén csak P-számok: a Q1 – P1 – Q2 egy ún. “triple”, ami jelentheti, hogy “Johann Sebastian Bach (Q1) fia (P1) Carl Philipp Emanuel Bach (Q2)”, de azt is, hogy “Az archívumban talált, XY raktári jelzetű levél (Q1) állítólagos feladási helye (P1) München (Q2).”

Q-számokat bármihez hozzárendelhetünk – emberekhez, dokumentumokhoz, eseményekhez, eszmékhez… Te döntöd el, milyen P-számokra van szükséged az általad kívánt állításokhoz. Az elemeket nem egy rögzített, módosíthatatlan kategóriarendszerben kell meghatároznod, a létrehozott állításaid pedig új árnyalatot és szilárdságot adnak az új vagy meglévő elemekhez. Ne aggódj, ha nem rögtön az első napon áll össze az adatmodelled. Hozd létre folyamatosan az állításokat, amikor csak szükséged van rájuk, közben figyeld, hogy érik el a kritikus tömeget, amellyel kiértékelhetővé válnak.

Minden állítás “minősíthető” – “Johann Sebastian Bach (Q1) felesége (P2) Maria Barbara Bach (Q2) házasság kezdete (P2) 1707. október 7. (dátum),  házasság vége (P3) 1720. július 5 körül (dátum).” Ezeket az állításokat ugyanakkor hivatkozásokkal is elláthatjuk: “erre bizonyíték (P4) XY egyházi évkönyv (Q3)”,”állítás forrása (P5) XYZ Bach-életrajz (Q4)”.

A rendszerben lehetséges egymással versengő értékeket megadni, mindössze külön-külön forrásmegjelölést kapnak, illetve rangsorolni is lehet őket.

Ilyen mélységben meghatározott triple-ekkel gyakorlatilag bármilyen állítást létrehozhatsz, ami viszont még fontosabb, ezzel lehetőséged nyílik állításokat létrehozni bármely nyelven. A rendszer Q- és P-számokkal működik, minden egyéb pedig címke, amit azon a nyelven adhatsz meg, amelyet fel szeretnél kínálni a felhasználónak. Ezen felül a szoftver automatikusan lefordítja a dátumokat és mértékegységeket az adott nyelv által használt formátumra. Ez a titka annak, hogy a Wikibase-platformokat mindenki a saját nyelvén szerkesztheti, miközben az egész világon olvasható szinte bármilyen nyelven.

Milyen segédeszközöket biztosít a szoftver?

Készíthetsz adatbázis-bejegyzéseket egyesével: nyisd meg a szerkeszteni kívánt elemet, menj a beviteli lap aljára, és kattints az “állítás hozzáadása”-linkre. Itt kell megadnod, milyen állítást szeretnél létrehozni. Nem szükséges fejből tudnod a P-számot, kezdd el begépelni a tulajdonság nevét a saját nyelveden, majd válassz a felkínált lehetőségek közül az automatikus befejezéshez. A platform tudni fogja az adott állítás P-számát. Az állítás második részét a következő szövegdobozban adhatod meg, szintén elég elkezdeni begépelni.

Excel- és CSV-listákból, automatikus bevitellel is készíthetsz adatbázis bejegyzéseket. (Itt találod a beviteli felületet, itt pedig egy rövid útmutatót hozzá.)

Az adatbázis-lekérdezéseket SPARQL nyelven kell megfogalmazni. Ez (sajnos) nem egy könnyű keresőnyelv, de végső soron annyira komplex, mint a futtatni kívánt keresések.

A SPARQL-t használók nem feltétlen tudnak SPARQL-forráskódot írni. Általában keresési mintákat tudsz használni, amik megmutatják, hol kell változtatnod a bevitt szövegen, hogy lefuttathasd a saját keresésed.

Amennyiben pontosan tudod, milyen típusú keresési lekérdezést kell futtatnia a felhasználóidnak, készíthetsz a könyvtárak megszokott online felületeihez hasonló, egyéni beviteli maszkokat, amelyek majd SPARQL-ben kommunikálnak az adatbázissal.

A szoftvercsomag tartalmaz illusztrációs lehetőségeket térképekhez, idővonalakhoz, hálózatokhoz, genealógiai kapcsolatokhoz, grafikonokhoz, stb. Nem kell letöltened egyéb, külső alkalmazásokat. A SPARQL-en keresztül kérheted az általad kívánt reprezentáció létrehozását. Gyönyörű bemutatót láthatsz vizualizációkból, ha felkeresed a Wikidata Scholia-projektjét.

Mit tegyek, ha a saját platformomon szeretném megjeleníteni az adatvizualizációm?

Ennek nincs technikai akadálya. Uwe Jung demonstrálta, hogyan használja az FH Potsdam felülete a Wikidatát adattárként úgy, hogy közben a felhasználók nem látják a háttérben lévő adatbázist.

Nincs semmi gond azzal, ha a FactGridet külső adattárként használod, és a saját kutatási projekted az egyetemed szerverén építed fel, ahol célzott adatbázis-hozzáférést teszel lehetővé saját keresősablonon keresztül.

A FactGrid CC0-licenc alatt teszi közzé az adatokat – ez azt jelenti, hogy lemondok a kutatásom jogairól?

Ha a Creative Commons 0-licencet választod, továbbra is teljes szabadsággal használhatod az adataidat, amire csak szeretnéd  – te irányítasz, és nem a kiadó vagy az adatokat kezelő platform. Ezen felül a CC0 azt jelenti, hogy az adataid szabadon felhasználhatóvá válnak mások által is. Mivel a közösség így bármikor kijavíthatja az észrevétlenül maradt hibákat, csökken annak a kockázata, hogy hosszabb távon elavuljon a kutatásod.

Néhány megfontolandó tényező: Tudósok számára első pillantásra a CC BY 4.0-licenc tűnik kedvezőnek. Ez engedélyezi az ingyenes felhasználást, amennyiben az megfelelően módon megjelöli a forrást. A gyakorlatban ez működhet szövegeknél (mint ez a blogposzt), mivel itt egyértelmű, hogy milyen hivatkozást szeretnénk látni: a nevünk megadásával, a publikáció címével, a kiadás helyével és dátumával. De szeretnéd, hogy az adataid idézetként szerepeljenek, például egy vizualizációban? Egy 1753 júniusában Párizsból Berlinbe küldött levél a térképen egy vonalként szerepel – hogyan lássuk el ezt megfelelő jegyzetekkel? Hogyan idézzenek téged, ha csak javításokat végeztél egy adathalmazon? Az “Így add tovább”-licencek még problematikusabbak: ezek az adatok szabadon hozzáférhetők bárki számára, amennyiben a további felhasználók is ugyanezekkel a feltételekkel osztják meg. Ez úgy hangzik, mint a szabad felhasználás melletti határozott kiállás. De egy al-felhasználó hogyan tudja biztosítani, hogy az ő al-felhasználói is betartják a licencbe foglalt feltételeket (főleg ha ez az al-felhasználó CC0 alatt teszi közzé az adatokat)? Az al-felhasználóknak általában azt tanácsolják, ne használjanak adatokat CC-BY vagy CC Így add tovább licenccel rendelkező platformokról.

A Wikidatával és a Német Nemzeti Könyvtárral közös vállalkozásunk egyetlen lehetőséget hagyott számunkra: hogy partnereinkhez hasonlóan szabadon felhasználhatóvá tegyük az adatainkat. A CC0-licenc által nem biztosított, hogy a további felhasználók is feltüntetik majd, ki gyűjtötte az adatokat, illetve felhasználásuk feltételeit.

A gyakorlatban a legtágabb nyílt licenc nem jelenti azt, hogy a FactGrid-adatok szerző nélküliek, épp ellenkezőleg. Mi azt szorgalmazzuk, hogy hivatkozzunk a kutatásra, és megelőlegezzük, hogy a Wikidata és a GND is boldogan feltünteti, ha a kutatás a mi platformunkról származik.

A FactGriden minden szerkesztéshez kapcsolva van a szerző neve. Ha egy kutatási projekt lényeges mértékben járult hozzá egy adatgyűjteményhez, akkor ezt jelezhetik egy külön jegyzetben, amelyet tovább lehet adni adatátvitelnél.

A Wikidatához vagy a GND-hez hasonló adatbázisok amúgy érdekeltek is a kutatások hivatkozásában – ez hozzájárul az adataik szilárdságához. A FactGrid abban a különleges helyzetben van, hogy mindkét szervezet számára olyan platformot szolgáltat, ahol a felhasználók olyasmiket csinálhatnak, ami saját, nagyobb platformjaikon nem engedett.

Mi történik, ha szeretném az adataimmal egy másik platformon folytatni a munkát?

Mivel szerzői jogi korlátozások nélkül vitted fel az adatokat, szabadon dolgozhatsz velük bárhol máshol. Valójában örülünk is, ha afféle inkubátor lehetünk kutatási adatok számára.

Mi történik, amikor FactGrid-felhasználók a “helyes” dátumról vitatkoznak?

A szoftver lehetővé teszi az egymásnak ellentmondó adatok kezelését – ez különösen fontos a történelmi kutatás területén, ahol gyakran találunk egymásnak ellentmondó forrásokat anélkül, hogy biztosan tudjuk, melyikük állítása igaz. A szoftverrel reprodukálhatjuk az ellentmondásos helyzetet, az állításokat pedig külön-külön alátámaszthatjuk hivatkozásokkal. Az eltérő állításokat súlyozhatjuk is egymáshoz képest – például a jelenleg irányadó állítást az egyéb variánsokkal szemben, vagy akár minősítőkkel az egyéni kiértékeléshez.

Tekintsük inkább érdekes helyzetként arra, amikor két kutató eltérő eredményekre jut. Sokkal rosszabb, amikor egy olyan platformon hibázol, ahol sosem lesznek kijavítva, és hitelteleníthetik az egész munkádat.

Miért kockáztassam meg az átláthatóságot rögtön a projektem kezdetétől?

Ez kemény dió, valószínűleg ez gátolja meg a legtöbb projektet, hogy használja a FactGrid erőforrásait. Az alternatíva egy platform, amihez csak a jelszóval rendelkező csapat férhet hozzá a projektet lezáró publikáció határidejéig. Így, szól az érv, semelyik versengő projekt sem tudja elcsaklizni a kutasi eredményeket. Senki sem látja, hol hibáztál az elején. Senki sem rögzíti, melyik adatot vitték fel asszisztensek és melyiket a projektvezető – ehhez hasonlók a feltételezett előnyei a nem átlátható munkának egy olyan platformon, amely csak a finanszírozás végével lesz online elérhető.

Az átlátható kutatás saját biztosítékokkal rendelkezik. Ha egy találsz egy minden eddigit felülíró dokumentumot vagy rögzítesz egy úttörő kapcsolódási pontot, akkor itt a lehetőség, hogy a saját nevedhez és projektedhez kösd az állítást. Ha holnap valaki ellátogat ugyanabba az archívumba és szintén felfedezi, amit te – pech, hiába. Te már rögzítetted a megfigyelést a platformon, amit a laptörténetben lekövethető változtatás minden kétséget kizáróan bizonyít.

Mindeközben a kollektív platform  meghívásként is működik az együttműködésre. Tedd egyértelművé a többi csapat számára, min dolgozol, hogy felvehessék veled a kapcsolatot.

Egy elméletileg biztonságos, csak a projekt végén nyilvánosságra hozott weboldal kockázatai komolyak. A felhasználókkal ekkor már nem lehetséges ötleteket cserélni. Az internetes jelenlét időzítése a projekt rohanós utolsó heteire esik, amikor már nem lehetséges semmiféle, koncepciót érintő változtatás. Ha a kutatást kizárólag egy könyves publikációhoz végeztétek, bizonytalan marad, mihez kezdjen a csapat a Word- és Excel-fájlokban összegyűjtött adatokkal. Senki sem tudja ekkor felvinni az adatokat egy nagyobb erőforrásba – egy ilyen késői fázisban a harmonizáció szinte megugorhatatlan akadály. Csak reménykedni lehet, hogy a könyv olvasói beszkennelik az összes lábjegyzetet, hogy a bennük lévő korrigálások elérhessék a könyvtári katalógusokat és a különféle Wikimédia-projekteket. A kockázatot itt a könyv jelenti, amely semmiféle hatással nincs a kollektív adatbázisra, illetve a digitális bölcsészet projektek, amelyek publikáció után elavulnak.

A jövő inkább egy újfajta hozzáállásban kell keresni egy közös, nyilvános adatbázis felé. A kutatóknak képesnek kell lenniük javítani és bővíteni ezt az adatbázist bárhol, bármikor hozzáférve. Az szükséges motivációt és biztonságot a kutató környezet jelenti, ahol megjelölhetik és idézhetővé tehetik saját munkájukat. Erre a Wikibase bármely más szoftvernél alkalmasabb.

Mi kell ahhoz, hogy a FactGrid befogadja a projektem?

A FactGrid-platformnak nincs láthatatlan mélyrétege. Bárki lekérdezhet az adatbázisból, és ugyanazt az eredményt fogja kapni akár be van jelentkezve, akár nincs. A személyes felhasználói fiók annyi előnnyel jár, hogy kiválaszthatod a kívánt nyelvet, miközben az adatokat böngészed, illetve lesz egy “szerkesztés”-link minden állítás alatt.

Ha szeretnéd betáplálni az adataid a FactGridbe, és ha szeretnél egy projektet futtatni a platformon, akkor szükséged lesz felhasználói fiókra. Ezt a valódi neved megadásával kaphatsz az adminisztrátoroktól. Ehhez az oldalon találsz egy “Request account” (felhasználó fiók igénylése) szövegű linket. E-mailben is felveheted velünk a kapcsolatot. Projektvezetők kaphatnak adminisztratív fiókokat, amivel kijelölhetnek csapattagokat, projekthez kapcsolódó személyeket.

Miután bejelentkeztél, felvihetsz adatokat nagy mennyiségben vagy végezhetsz meghatározott javításokat bármelyik elemen. Minden változtatásod a felhasználói fiókodhoz lesz kapcsolva. Mások visszavonhatják a szerkesztéseid, de nem nyomtalanul, dokumentálva lesz az elem történetében, mindenki láthatja.

Ha egy összetettebb projekten szeretnél dolgozni, —

  • ami lehet személyes családkutatás,
  • lehet egy egyszeri vizualizáció egy szemináriumi dolgozathoz,
  • vagy akár több ezer tételnyi adat bevitele egy 5 éves projekt folyamán

— egyeztess a többi felhasználóval és a platform szervezőivel. Nem (feltétlen) fogunk egy nyilvános egyetértési nyilatkozatot aláírni, de a blogunkon hírt adhatunk a projektedről, hogy eljusson mindenkihez a platformon. A munka akkor válik igazán izgalmassá, amikor mások befejezett munkáját módosítod, illetve amikor más projektek résztvevőit inspirálod az általad bevezetett modellezés használatára. Nem kötelező átbeszélni az adatmodelleket a többiekkel, de a modellek megosztása segíthet a kutatásodnak új embereket elérni, illetve felhasználhatók lesznek mások által létrehozott lekérdezésekben vagy vizualizációkban.

A szoftvert arra tervezték, hogy kezelni tudja mind az olyan állításokat, amelyek csak számodra érdekesek, mint azokat, amelyek az eredeti kutatási témádnál jóval távolabbra elérnek majd.

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).
Jack Kirby, “The Fourth Dimension is a many splattered thing!”, Alarming Tales, 1 (1957. szeptember).

FactGrid FAQ – Why should I use FactGrid for my research project?

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).

auf Deutsch
en français
magyarul

What is FactGrid?

FactGrid is a Wikibase installation — that is is both, a regular wiki and a database which you can use to make statements about objects of your interest — statements which you can then handle in practically any language in big data sets.

The platform is run by the Gotha Research Centre and hosted by the ThULB Jena. It addresses projects with a specific interest in historical data and is part of the German National Research Data Infrastructure NFDI4Memory.

In joint ventures with Wikimedia Germany and the German National Library’s GND we are trying to bring this platform into the upcoming consort of federated Wikibase instances as a resource for research data.

Why should I use FactGrid for my own research?

The biggest argument for a FactGrid account is the unbeatably flexible software, Wikibase, which we have managed to implement in a pilot project with the help of Wikimedia Germany – outside its primary location, Wikidata:

  • You are looking for a software that speaks practically any language — a platform on which you can enter data in your language and allow others to read your data in their languages? Wikibase is this software.
  • You are looking for software in which you can transparently coordinate a whole team? In Wikibase this is as easy as in the Wikipedia software MediaWiki .
  • You are looking for a database software that can do everything Digital Humanities databases normally want to do: network analysis, map representations, complex linked searches, timeline representations (in various formats) – a software that almost acts like human language yet provides full database services? Wikibase is this software.
  • You have data from previous projects which you want to build on? Wikibase has large-scale automatic input options.
  • You want to make sure that that other projects will actually use your data? Use a platform that allows the download and further work with your data offline in Excel or online in ever new projects.
  • You want to ask entirely new questions in your research? In Wikibase you can link any sort of objects with any kind of statements of your interest.
  • You are worried what will happen to your data and presentations once your funding is over? Work on a platform where you do not stay alone and where, thanks to the CC0 license you are using, you encourage colleagues to continue right were you stopped!

If you are looking for a long term perspective this is what we are trying to offer in our present joint venture with the German National Library. We will base our platform on GND data in order to serve as a broad public tool and with the aim to become a player in the emerging landscape of “Federated Wikibase installations”.

Why not use Wikidata right away?

That is a legitimate question to ask. There will be projects (projects that are mainly using data) for which Wikidata will be the better platform. The FH Potsdam’s “Archivführer zur deutschen Kolonialzeit” has demonstrated the beauty of working directly on Wikidata; we talked about this with Uwe Jung, who demonstrated the technical solutions they have chosen in Potsdam.

On the other hand, there remain basically two things which you will not be able to do on Wikidata or on a platform like the GND: Wikimedia projects (and the GND) have strict “No original Research” policies and observe fundamental decisions to operate on “criteria of notability“, which will not allow the arbitrary opening of database objects and the innovative object relationships researchers would like to test.

Wikidata and the GND focus on information that has already been published and on non-research workers who feed the respective databases from published research. You will not be allowed to state a “working hypotheses” of “your research” on these platforms. You will not be able to create entities with the aim to run “nothing but a statistical analysis” on them at a far later stage of your work.

In FactGrid we encourage the use of the platform as a heuristic research tool.

  • Create database objects on the platform, no matter what their relevance in an encyclopedia or in library catalog could be.
  • Risk provisional chronologies as working hypotheses along with your personal assumptions.
  • Use FactGrid in order to make unconventional statements that are presently interesting only in your research project – the software gives you this freedom.
  • Create specific database objects that state your research in all the data sets which you have substantially modified and become able to submit your research with the particular item as the envelope to your funding institution.
  • Risk any new thesis on the platform and state your respective view with a database object number as a “micro-publication” in order to claim your ingenuity on the data base.

FactGrid is free – how does it work?

The software is freely available and under development in the larger community of Wikimedia projects and among the institutions which are going to use Wikibase over the next years.

The FactGrid platform is presented by the Gotha Research Centre on a virtual server of Jena’s University Library free of charge.

All the Wikidata-tools are available to our users. These include all the standard applications of regular Digital Humanities projects.

With both software and tools being open source you can use any favorite software company you are working with, to generate the specific application which you feel you need.

If you feed your tools into the open compound this will be your best way to make sure that future projects will continue their development.

If you aim at technical solutions which you want to sell with financial profit that again will not be restricted by the software license. You can freely commercialize whatever you create on the basis of the open software.

What do I do with unorthodox research interests?

Wikibase is groundbreaking in its data modelling. Essentially you are only creating relations between Q-numbers (or relations between Q-numbers and dates, Q-numbers and space coordinates, Q-numbers and media files, Q-numbers and URLs).

The software does not know what kind of relationships you are stating – these again are just P-numbers: Q1 – P1 – Q2 is a “triple” and can mean “Johann Sebastian Bach (Q1) is the father of (P1) Carl Philipp Emanuel Bach (Q2)”; it can mean just as well “This letter which I found in the archive with the shelf mark XYZ (Q1) has allegedly been sent from (P1) Munich (Q2)”

Q-numbers can be assigned to anything imaginable – people, documents, events, ideas… You decide what kinds of P-numbers you need in order to make statements of your interest. You do not define objects in a system of fixed categories; your statements are adding colour and solidity to whatever object you create as you go along. Do not worry if you do not have the data model on day one. Make statements when you suddenly want to make them and see how they gain the critical mass that can eventually be evaluated.

All statements can be “qualified” – “Johann Sebastian Bach (Q1) was married to (P2) Maria Barbara Bach (Q2) beginning on (P2) October 7, 1707 (date) ending (P3) about July 5 1720 (date).” All of these statements can in turn be equipped with references : “this is clear from (P4) the church book of… (Q3)”,”this is stated in (P5) the well known Bach biography XYZ (Q4)”.

The system allows competing claims at any time. They are simply introduced with their different sources and can be balanced against each other.

You can create practically any normal language statement with triples of this depth of specificity; but above all, this opens the door to the world of statements in all the various languages you might want to speak: The system operates with Q- and P-numbers; the rest is labels in languages which you want to offer to your users. The software will in addition translate dates and quantities into other formats; this is basically the secret that makes it possible for Wikibase platforms to be edited by people in their own languages and to be read by the world in practically any other language.

Which tools does the software provide?

Database entries can be made one by one: Open the object-id in question, go to the bottom of the input page and click the “add statement” link. You will now be asked for the statement you want to make. You do not need to know the P-number. State the property in the language you are using and click at the auto complete you are eventually being offered. The platform will now use the P-number of that statement for you. Complete in the next box that opens your statement. You will again get suggestions to use as you are typing.

Database entries can also be created and substantiated in automated inputs from Excel or CSV lists. (This is the input mask and this is the short guide to it.)

Database queries have to be formulated as “SPARQL” queries, a search language that is (unfortunately) not that easy to use, but that is eventually as complex as the searches you might want to perform.

SPARQL users do not necessarily know how to write SPARQL source code. You usually use sample queries that tell you where you have to change the input in order to run your particular search.

If you know exactly which type of search queries your users should run you can create your own input masks just as you know them from conventional online library interfaces, which will then speak SPARQL with the database.

The software package includes illustrations on maps, timelines, networks, genealogical relationships, graphs, and so on. You do not need to download particular applications. You will ask SPARQL to produce the representation you are trying to get. The Scholia project on Wikidata has a beautiful first presentation of some of the visualisations.

What do I do if I want to give my very own data representations on my own platform?

That should not be a technical problem. Uwe Jung demonstrated how the FH Potsdam interface uses Wikidata as its data repository, without letting users see the database they are accessing.

There is nothing wrong with using FactGrid as an external repository and building up your own research project on the server of your home university, where you can offer targeted database accesses under a typical search template of your choice.

FactGrid basically licenses data to CC0 – does that not mean that I give up all rights to my research?

Opting for the Creative Commons 0 license means essentially that you continue to be free to do whatever you want with your data – you, and not your publisher or the platform that received your data under a scheme, continue to control. But above all, the CC0 license means that your data becomes freely usable and that you can thus reduce the danger of obsolete research in the longer run – others will continue to root out the ugly mistakes you could not hope to correct.

Some basic considerations: CC BY 4.0 is at first glance the license that scientists will prefer: It allows the free further use as long as it receives the accurate citation. In practice, this will work for texts (such as this blog post); here it is clear how one would like to see the text cited: with a reference to one’s own name, with the title of the publication, the place of publication and the date. But do you want your data quoted let us say in a visualisation? A letter sent from Paris to Berlin in June 1753 will be a line on a map and how should this line be properly annotated? How do you want to be quoted if you only improved a data set? “Share alike” licenses are even more problematic: “These data are freely available if the subsequent users keep it just as freely available.” That sounds like the ultimate plea for free use. But how can a sub-user ensure that his sub-users, in turn, will respect your license agreement (especially if this sub-user is offering his data on CC0)? Sub-users are well advised not to use data from CC-BY or CC Share-Alike platforms.

Our joint ventures with Wikidata and the German National Library left us only only one option: to make our data as freely available as our partners: CC0, that is without ensuring that subsequent users will still specify exactly who collected the data, and what third-party users are allowed to do with that data.


In practice, the maximum open license does not mean that FactGrid data is data without authorship, quite the contrary. We suggest that research is cited and assume that Wikidata and the GND are only too happy to state research from our platform. All changes are linked to respective the author names. If research projects have worked more substantially on a data set, they will have stated this in a separate note on the data set that can now be adopted with the data transfer.

Databases like Wikidata or DNB’s GND are in fact interested to quote research – it boosts their data solidity, and FactGrid is here in the unique position to give both institutions a platform on which people can do what they cannot do on the respective larger platforms.

What happens if I want to continue working with my data on another platform?

Since you entered your data without a copyright restriction, you are free work with them on any other project of your interest. We actually like to be “just an incubator” for research data.

What happens when FactGrid users argue about a “correct” date?

The software makes it possible to handle contradictory data. This is particularly interesting in the field of historical research, where we often have conflicting documentary evidence without being able to determine the correct statement after so much time. The software makes it possible to reproduce such a contradictory situation. Numerous statements can be given side by side with their various respective sources. You can then still balance the statements against each other – either by turning one of them into the statement to privilege in future searches and/or by adding qualifying statements with your personal evaluations.

If two researchers come to different results, think of it as the situation you should actually be interested in. It is far worse that you have made mistakes on a platform where they will never be corrected and where they eventually discredit all your work as obsolete beyond repair.

Why should I risk transparency in my project right from the start?

This is likely to be the toughest issue that currently prevents projects from using the resource which we have opened. The alternative is the resource, which is accessible to the team only under passwords until the publication deadline is reached almost at the end of the project. No competing project can snatch away findings, so the theory. No one sees where you have initially made a mistake. Nobody records what assistants are typing in and where the project leader is involved – such are the presumed advantages of non-transparent working on a platform that will only go online at the end of your funding.

Transparent research offers its own securities: If you find a groundbreaking document and establish a decisive connection, then this is your chance to fix the statement to your name and project. If someone makes the same discovery tomorrow in the archive you have just visited, tough luck. You will have recorded your observation with a link in the version history which your rivals will not be able to deny.

At the same time, the collective platform expresses the invitation to cooperate. Make it clear to other teams what you are working on and allow them to contact you on the platform!

The risks of the allegedly secure website, which is only going online at the end of the project’s funding, are serious: The time for an exchange with users is over. The internet presence goes online in the heated final weeks while the project is totally unable to react with more conceptual changes. If you have done research solely for a book publication, it will remain unclear what you and your team should do with the data, you have still collected in Word files, and Excel spreadsheets. Nobody will be able to feed all these data into any resource – the harmonisation at this late stage will be an insurmountable obstacle. You can only hope that readers of your book will scan all your footnotes for corrections that should reach our library catalogs and the various Wikipedia projects. The risk is here the book that has no influence on the collective data base and of DH projects that become obsolete right after their publication.

The future should lie in a new attitude toward the public data base. Researchers should be able to correct and to further widen this base wherever they access it. The incentive and the security they will need here is the research environment in which they can mark their work and make it citable. For this Wikibase is better equipped than any other software.

How do I get my project accepted on FactGrid?

The FactGrid platform has no invisible deeper layer. Anyone can query the database and the queries will give the same information whether you are logged in or not. Your personal user account just has the advantage that you can now switch to your favorite language when looking at data, and that you see the edit link on each statement.

If you want to feed your own data into the platform and if you want to run a project on the the platform, you will need an account. These are given under real names by the administrators. The software provides an “account request” link. You can also contact us via email. Project leaders can receive administrative accounts which they use to assign to team members and users of their interest.

Once logged in, you can enter data in bulk or make specific corrections wherever you feel like. Any input will be connected to your user account. Others can undo your edits but not without leaving a documented mark of that intrusion in the version history – visible to all the world.

If you want to work on a more complex project —

  • that can be personal family research,
  • it can be a single visualization you need in a seminar paper,
  • it may just as well be the input of thousands of records in a 5-year research project

— speak with those on board and those organising the platform. We will not (necessarily) be interested to sign a public memorandum of understanding with you but it might be cool to advertise your project on the blog and to make it known on the entire platform. Your work becomes exciting, where you modify the work others have already done and where you encourages players from other projects to adopt good models which you are introducing. You do not need to discuss data models with all the others but using models with all the others is also a way to spread your work and to make it appear in queries composed by others, and visualizations you did not think of.

The software is designed to manage both: unique statements which only you are interested in and statements that will spread far beyond your own initial research interest as you are now feeding the unforeseen queries which others will run on their and your data.

Jack Kirby, "The Fourth Dimension is a many splattered thing!" from Alarming Tales, 1 (September 1957).
Jack Kirby, “The Fourth Dimension is a many splattered thing!” from Alarming Tales, 1 (September 1957).

Memorandum of Understanding between the University of Erfurt and the German National Library – to base the FactGrid on GND data in a joint project

German Version

We are proud to announce a new and massive Wikibase project that should keep a large community busy for far more than a year: Last month the president of the University of Erfurt, Prof. Dr. Walter Bauer-Wabnegg, and Dr. Elisabeth Niggemann, director-general of the German National Library in Frankfurt and Leipzig (DNB) signed a memorandum of understanding that aims to bring GND data into the FactGrid – on a grand scale.

The GND, the German Integrated Authority File, is an authority file of millions of persons plus corporate bodies, conferences and events, geographic information, topics and works – designed to shape the exchange between libraries, archives and academic projects in the DACH countries of Germany, Austria and Switzerland.

integrating the GND into the FactGrid had been our constant topic of discussion during the last year. A Wikibase instance becomes a cool thing to contribute to, as soon as it becomes the research tool that you would use yourself in your research. GND data links into the world of open data; they clarify who or what you are speaking of in your research in all German-language contexts – and they will reach out to the other global authority files and to the universe of library data.

In April 2018 it became clearer that the FactGrid would eventually be one of several Wikibase instances which could and should in this case aim for a larger federation. Early in June it transpired that the German National Library was on its way to test Wikibase in a software evaluation, with the aim to run possibly about ten Wikibase instances in a constant exchange with each other. That was when we contacted the DNB with our own agenda to import their data. We wanted to try, so that our proposal, could become a platform for “original research” – a platform without GND or Wikidata criteria of notability – in the evolving network of Wikibase platforms. Users will be allowed to create Q-Numbers for infants who died right after birth on FactGrid, and the GND and Wikidata will be free to decide under their criteria of notability and relevance, whether they would like to use our information – information they can now quote as original research from the FactGrid platform (with the detailed information of the projects behind this research).

Whilst the GND is CCO and free to be copied, the open joint venture with the German National Library aims to bring transparency into the data input. The more transparency we can bring into all the design decisions in this early stage, the better the wikibase platforms we are heading towards, will eventually be able to communicate with each other.

Now a team has to be formed. The German National Library and the Gotha research institutions of the University of Erfurt will send members into the team. The question is: Will we be able to broaden this team? We should have experts from the Wikimedia communities on board – people who know Wikibase and Wikidata, people who are used to community work on a regular wiki.

  • We would like to attract people who know how to formulate SPARQL searches and who will be able to test data models and make suggestions for the improved data models we should use, in order to handle the massive data sets we are expecting.
  • We’re looking for Wikibase experts who know how to bring in tens of millions of records into a Wikibase installation, and who know how to interconnect these records with genealogical and geographical links.
  • We do not yet know how we will keep the FactGrid manageable with respect to the wave of doublets and name parallels we are facing: The GND has these name parallels in unprecedented numbers. We will have to find ways to quickly inform researchers whether a person they have found in a document is already on the FactGrid or whether they will have to create the item. The hunt for items to be merged will become a permanent issue and we do not yet know how to technically support a community on this collective quest.
  • We will create new and complex fields of expertise: Millions of personal data sets will come with career statements. The FactGrid will turn all these statements into Q-Items, which we will have to organise in order to allow sociological searches for instance. The FactGrid project on historical jobs and their evolution will be only one of these projects.
  • We need players with Wikipedia experience: Though we will restrict ourselves to clear name accounts, we widely invite users with professional to private ambition to join the platform with their projects – whether they are focused on private genealogy or on publicly funded historical research.
  • We will have to provide a simplified FactGrid user interface that will bypass the SPARQL QueryService and the mushrooming Wikibase input pages. Magnus Manske’s Reasonator might become our standard interface for regular users, who will access the FactGrid as if they are accessing library catalogues – through organsied input forms.
  • We will eventually need help with database maintenance. It is particularly unfortunate that our project is primarily the work of historians, who do not always have a keen eye on how to optimally supply this technology.

The FactGrid will grow – and it will offer plenty of space for people to develop their own projects within this growth.


Scan of the Memorandum of Understanding (in German)


More

  • Barbara Fischer & Jens Ohlig, “Neues Testfeld für Wikibase: Eine Bundesbehörde geht auf Expedition im Wikiversum.” 2019-05-09 at https://blog.wikimedia.de