Needed thing #4: A module to state original claims (and published research)

The Problem

Original research means that we will (also) have to deal with statements that have not been published before. So far this is a huge problem for any researcher. Should she make a claim that was never made before – minutes after she found the archival record to substantiate the spectacular claim? You better wait until your book is out – which can take a couple of years, and if you still need a database to do your research you better work on a platform where your work is invisible until then.

The platform with immediate visibility of your work is at the same moment a massive advantage: If you publish the observation minutes after you made it, you will have made your claim and you can from now onwards refer to it. That, however, means that FactGrid claim has to be made publicly, visibly connected to your research, your name, with a specific URL that comes with a publication date.

The FactGrid must be able to turn any statement which is made on the database into a micro publication. You make the claim and you give the source with all the information about you including your evaluation and the details which any future research should continue to offer. The Wikibase Interface of our dreams should offer footnotes on each claim, every note nicely wrapped up for anyone to grab and to repeat in his own texts.

The more complex source attribution will have more advantages: It will allow researchers to fill the database with hypothetical statements. These will be marked as such and enter the test run, for you will now be able to see whether a hypothetical date (for instance) of a letter fuses into the data environment you are creating. You can immediately work with colleagues on a premise where you feared them as rivals who could steal your information.

Model solution

The FactGrid source attribution will have four sections. Users should be guided with drop down menus where possible. We generate a new Q-item to quote in the end:

Section 1: Published elsewhere or original research? (pull down menu plus input fields)

  1. This statement is already publicly circulating. (Input fields:) Q-number of the publication (plus field for more specific reference like a specific page number).
  2. This is original research to be credited as such. (Input fields:) Q-numbers of the researchers or team to be credited (plus date stamp and url to quote the entire module).

Section 2: the evidence

  1. Q-number(s) of the piece(s) of evidence (plus field(s) for detailed reference like page number(s)).

Section 4: evaluation (qualifier to the previous via pull down menu)

  • The claim can be taken for granted with the evidence given.
  • The claim is based on additional conclusions (stated in section 4).
  • No evidence given, yet the claim is generally accepted as fact
  • The claim is obsolete (for reasons discussed in section 4).
  • The claim is/was hypothetical (the assumptions are stated in section 4).
  • The claim is valid within the fictional universe.
  • The claim has a propaganda value.
  • The claim is part of a religious creed (see the discussion of section 4).
  • The claim is personal/family knowledge.

Section 4: discussion (link to the statement’s discussion page)

Use

The source statement will ideally contain all the information needed to (automatically) generate a footnote which can then be used by the Reasonator the FactGrid’s equivalent (see our needed thing #3), in any Wikipedia article or in any other publication referring to the claim.


Published also here: https://www.wikidata.org/wiki/Wikidata:FactGrid/Needed_thing_No._4:_A_module_to_state_original_claims_(and_published_research)

Needed thing #3: An attractive Interface for browsing and reading Wikibase information

The Wikibase software has been designed to serve underneath the +200 Wikipedia installations, it is offering its services in SPARQL-queries but it does not aim at people interested in the facts collected on an item of knowledge.

Magnus Manske’s Reasonator is the tool which turns Wikidata information almost into articles – in any language. The page on Q13339, Johann Sebastian Bach is, as it turns out, in many ways superior to the 200+ competing Wikipedia articles on Bach: It has one sinle source to be edited by users world wide. It shows at a single view what it has to offer – you do not crawl through well balanced sentences, which might not at all offer the information you are looking for.

But the Reasonator has its fundamental drawbacks: Technically you are on a platform that uses Wikidata information – not on the global Wikidata interface. Practically and organisation-wise you are on extraterritorial space when it comes to future developments. The Reasonator is Magnus Manske’s dream child. It is not part of the package Wikimedia will develop as the universal Wikidata front-end (because any such front-end would immediately rival the 200+ Wikipedias?)

The following thoughts aim at an “Interface” one would like to have with any Wikibase installation on whose and what technology whatsoever:

What the global “Wikibase Interface” should be able to do (and what it should avoid)

  1. Pages on items of knowledge (i.e. on Q-numbers of the installation) should not rival the written article (with automatically generated language statements).
  2. The interface should focus on the presentation of all the facts on a specific question. Get the first three entries of the list and get the complete list only if you click at more. Use the interface to get all the letters Leibniz has written, all the works composed by Bach, all the people Luther is known to have met plus dates and locations.
  3. An edit option leads from the specific statement on the Interface page to the specific Wikibase input section that is generating the statement.
  4. Users who are reading a biographical Interface page can press “edit via form” and they will be led to an input form for biographies with subsections to open. This is particularly useful on any page with fragmented and sparse information, since Interface readers will not necessarily have a clue what a Wikidata property is, and where to find it. They need inspiration of what questions they possibly could answer. See our Needed thing # 1: The technical solution that enables researchers to create input forms for the specific requirements.
  5. Any statement on the Interface page is referenced on page in a footnote (see Needed Thing #4: A module to state original claims (and published research)) so that users can grab the footnote and get it into the Wikipedia they are writing or into the research paper or book under their hands.
  6. The interface can present media and extended texts. A page on an archival document or 18th-century book must be able to offer the scans and a searchable text transcript (users who detect transcription mistakes must be able to correct the mistakes on the spot, through the interface). See one of our Illuminati-document pages for the requirement to be met.

Magnus Manske’s Reasonator is the Wikidata exploit that has taken the step into the data-driven alternative to Wikipedia articles. We should see the advantages: We leave the world of tediously constructed texts and all the confrontations these texts are bound to sparkle between want to be authors and offended readers. We get information that is actually generated in a global effort – where Wikipedia has been generating national communities so far with all their massive problems. We can aim at complete collections of facts. Do not press for “more” on a subject if you do not want to get the names of all the children Johann Sebastian Bach had – but use this source if that is what you want to know. We leave the debates of the various “notability” wars we are presently leading in or 200+ Wikipedias – the debates on what a respective “community” feels people should know, and what they feel one should not necessarily be bothered with.

We must reach the point where we see that Wikidata has actually merits of its own as a new additional source in the Wikimedia universe – and this is what Magnus Manske’s Reasonator has been doing almost in the shadow so far.

Links & More


Published also here: https://www.wikidata.org/wiki/Wikidata:FactGrid/Needed_thing_No._3:_An_attractive_Interface_for_browsing_and_reading_Wikibase_information

Needed thing #2: A logo and our own design

The facts all contribute only to setting the problem, not to its solution.
        Ludwig Wittgenstein, Tractatus Logico Philosophicus 6.4321

The FactGrid still needs its own cohesive design. The name is a modest allusion to Wittgenstein’s Tractatus and his idea that we see the world through a grid of factual statements. It was not that difficult to correlate this thought with images – looking backwards and a across cultural borders. The blog’s main page uses these changing images with humour and as inspiration.

That, however, is all we have at the moment – leaving a lot to be done. The different software platforms – our blog (WordPress), the Wikibase installation (Wikimedia design), and Magnus’ Manske’s Reasonator child do not really go together design-wise.

  • The project does not have a logo.
  • We are presently using Corbel on the FactGrid’s blog, a font with space to breathe, modern with its sans serif design and yet conservative with its medieval numerals and ligatures… is there an open source alternative? And: do we need to go open source with the font?
  • The database still has the Wikidata design (and basically the design of all the central Wikipedia projects). The blog is more in the direction to go. The database should, however, stay in close contact with its wikibase mother, so that anyone working primarily on the mother project can immediately feel right at home on our platform.
  • The FactGrid’s Reasonator interface should enjoy greater freedom to adopt a unified design since we are here mostly interested in an interface that represents information. We should here go for a design that is open to bigger representations of maps, images, models of objects, since our projects will be forced to produce show cases.

These remarks are will not yet serve as a specified task book, they should rather set a direction.


Published also here: https://www.wikidata.org/wiki/Wikidata:FactGrid/Needed_thing_No._2:_A_logo_and_our_own_design

Needed thing # 1: The technical solution that enables researchers to create input forms

Wikidata’s Wikibase installation has been filled almost entirely in massive automated data inputs. That is probably why input forms were not exactly the first priority.

Our database will focus on researchers and regular users whose tasks will call for modules which they can get used to. The historian might sit in an archive with the task to register some 200 documents of a law case. The documents have to be dated, information about authors, the institutions, and addressees has to given on each document. The private user might want to give biographical information about a distant family member with the aim to augment his family’s genealogy. Both are used to input forms. They will never have heard of “triples”, their ideas of “properties” will be inappropriate, they will not be able to use complex Excel-commands in order to prepare an input via QuickStatements.

Requirements

What we need is a technology which enables projects to create their own input forms:

  1. Research projects must be able to define and modify such input forms – using the properties they have created or found on the database.
  2. It should be possible to define and explain the particular input field – whether this field calls for an item (with a Q-number), a date or a numeric value etc. Predefined pull down menus will be particularly useful in a lot of cases.
  3. The ideal input form will give indications whether an item is already in the database by auto-completion and through suggestions.
  4. The tool should be able to create database items with new Q-numbers on a first input.
  5. It must be possible to return to a form and an item of interest once fresh information can be added so that bigger teams or a crowd can work in successive sessions on the same items. The Q-number could be the entry point.
  6. We should be able to nest forms, that is to include specific modularised forms in a bigger form: A biographical input will open with basic questions and it will then offer specific modules on the genealogy, places the person has lived and visited, education, degrees, memberships or works. A membership module for the Illuminati will differ from a membership module for the British House of Commons since being a member will raise altogether different questions in either case. The option of specific modules is necessary since we might get rare but complex options of interest to specialists only and since we should be able to duplicate entire modules: If a person is employed by different companies we get the same questions open again: From when to when? Which company? Where stationed? What position? What salary?

Use cases

Biographies will be the most interesting test field. Most users will have augmented their own CVs with biographical information more than once in their lives.

The document description will be the most interesting input form for historians to use – and a use case of its own practical value. We would test here the use of the database at the entry point where knowledge is produced with a tool that should be more handy than the usual individual word files which researchers are using for excerpts and random bits of information. The FactGrid document description could be used by archives in turn to gain the metadata users usually generate for their own purposes.

Status

Erfurt University funded a prototype development (see: https://database.factgrid.de/wiki/Web_Forms). We became able to generate input forms on the platform, smoothly using any properties a research team would gather on a specific module. It turned out to be more difficult to access such a form again at a later stage (as described in requirement 5 above).


Published also here: https://www.wikidata.org/wiki/Wikidata:FactGrid/Needed_thing_No._1:_The_technical_solution_that_enables_researchers_to_create_input_forms