January 13, 2020

Hindi as a world language

A map of the Wikimedia foundation shows the readership of the famous encyclopedia for each country in the world, https://stats.wikimedia.org/wikimedia/animations/wivivi/wivivi.html The Hindi language has a wide distribution all over the world. It is used by India to retrieve information, but the language is also spoken in the US, China. The overall population which speaks hindi is 560 million according to the info box, but in reality it's much more. It's only a conservative estimation and the number is growing. Hindi has a good chance to become the worlds famous language. In india alone, over 50 million pageviews each month are generated by Hindi speaking users. That means, they are using the WIkipedia in their mother toque to retrieve and write information about their own country and the world in general. The overall population in India is 1300 million which means, that in the future the amount of pageviews in the Hindi language will grow.

January 10, 2020

The german wikipedia debates how to find new authors ...

Under the URL https://de.wikipedia.org/wiki/Wikipedia:Kurier the German version of the signpost informs the Wikipedia community and the public as well about the project status. The main topic in January 2020 is, that the amount of active contributors is to low, and the community argues about the reasons. The idea is to increase the amount of authors, but nobody knows how to do this exactly.

Even if my German language skills are great, it's hard to follow the debate. Not because of the vocabulary but because i have the opposite fear. The problem with Wikipedia is, that the amount of authors is too high and could explode in the future. What does that mean? On the first impression, the Wikipedia project is protected against chatbots, because a chat bot is not able to create a valuable edit. This is true for a complex article which contains lots of natural language and is equipped with extensive references at the end. The danger in Wikipedia is, that chatbots are utilized to generate a certain sort of Wikipedia content which is highly structure. This is called a stub. A stub is small article which contains of two sentences and can be generated from a RDF-ontology by a knowledge-to-text system.

The resulting stub will read like a normal article, except that it was generated by a computer program. In contrast to humans, it's very easy to make a copy of the computer program. A simple unix command like “cp chatbot.py chatbot2.py” is enough to create as much wikipedia authors are needed. The funny thing is, that according to the published papers at Google scholar, simple biography stubs are generated with bots already. That means, the amount of Wikipedia authors is higher than the german chapter is aware of it.

January 06, 2020

The pros and cons of Wikiprojects

An interesting meta-section in the Wikipedia encyclopedia is a Wikiproject. In contrast to the portals, most wikiprojects are available today. That means, it's a barrier to put a WIkiproject into the deletion discussion. On the other hand, the propability is high that this will happen in the future, because a WIkiproject has the same problem like a portal: low amount of users, and low benefit for the overall Wikipedia.

Instead of arguing pro WIkiproject deletion, it make sense to use the time to hear what the experienced users from Wikiproject have to explain why the project make sense. The interesting point is, that Wikiprojects are the key for reaching an audience in the university. The Wikiproject medicine for example was used sometimes by medical students in courses. It's a low entry option to become familiar with Wikipedia.

Let us investigate what medical students are interested in: they are not motivated to learn something about physics nor computer science, but they stay within their own subject. If somebody studies medicine he likes to read books about it, and that means, only books about medicine. So it's a natural choice to create a subpage in Wikipedia to coordinate a team of medical students who are interested in improving existing articles. That is basically spoken the idea behind Wikiproject and the reason why they were founded in the past.

Suppose the idea is to delete all wikiprojects, what is the future participations of medical students in Wikipedia? On the first one a deletion, would block all efforts to contribute to Wikipedia especially if the own domain is focussed on a single subject. Perhaps it make sense to go a step backward and describe what science in general is about. The idea behind science is to become a specialist on a single subject. The idea is to reduce the scope. The question is, where is the right place in the Wikipedia project for doing so?

The answer is very simple: Reducing the scope is done with keywords. A keyword like “Immune System” is more specialized than the general term “Medicine”. Contributing to the Wikipedia project is possible with article request, maintenance request, deletion request, peer review request, and photo request. That means, the user has to open the page and put in the term into that page. For example, the article request page contains of all the domains, like physics, literature and medicine. What the user is allowed to do is to put his specialized keyword into the section “medicine” in the article request page.

This sounds a bit complicated but it's equal to tag a request. That means, all the request are handled as request but they are tagged with domains like physics, medicine and literature. This kind of interaction provides the same feature like a wikiproject but is adressed to a broader audience. The advantage is, that all the users have to observe the article request page. That means, even non medical experts are allowed to enter new article requests.

The wikiproject concept is equal to a decentralized project coordination. The idea is to split the Wikipedi a into sub-sections which are domain specific. This produces a lot of inefficiency. The better is to centralize the requests and tag a request with it's domain. That means, if a user likes to improve the project, there is only one page available which is the request directory. And all the issues are put into this single page.

The effect is, that the page view is higher, more people will monitor this page and the productivity is better. Wikipedia is on a path towards this goal. Since mid 2019 many portals are deleted already and in the future, the wikiprojects will follow. The switch from decentralized requests for writing new articles into a centralized request page will make Wikipedia more professional.

Let us imagine how a medical student can contribute to Wikipedia. What the user has to do is to find a very complicated medical term which has no WIkipedia article right now. This complicated term is put into the article request page, because Wikipedia should be explored into that direction. The same user can write the article for the term, and if he is done, he deletes the term from the list. The trick is to put only terms on the list, which are specialized. That means, it is not directed towards a mainstream audience but the amount of papers about the term is very low. The resulting article will attract very few readers. At the same time the reputation for creating such an article is high. What researchers are doing is to aggregate knowledge about a complicated seldom used term from a specialized domain.

Let us take a look into the reality to determine if the proposed workflow make sense. https://en.wikipedia.org/wiki/Wikipedia:Requested_articles/Medicine is the request page for all medical terms. It has a pageview of 2 per day. Which is very low. It provides a links to other languages like English. And it has some keywords like “Marshall protocol” which is a specialized subject within immune system diagnosis. According to the changelog the page isn't edited very often.

In contrast, the Wikiproject Medicine has a dailypage view of 86. And in the talk section, lots of domain specific discussion is available. It seems, that today, the decentralized wikiproject medicine is more attractive to the users, than the centralized version. The problem with the Wikiproject medicine is, that even this specialized portal doesn't fulfill the needs of the users. So they have created many subprojects: Wikiproject Anatomy, Wikiproject Physiology and so on. THe result is some kind of WIkiproject spam, in which the amount of projects is growing, but the team behind it, is doing nothing. On the other hand, i think the Wikiproject idea is a good possibility to learn because it shows, what the users are interested in. In most cases the idea is to specialize on a single subject which is a good idea because this will bring science forward. Let us click on the item Wikiproject anatomy and observe what comes next. In the anatomy section a new subfolder waits for the user, it's called Category:Anatomy articles by topic, that means, the user can decide if likes to read articles about subsection of anatomy.

January 05, 2020

Wikipedia is restructuring it's portals

https://en.wikipedia.org/wiki/Wikipedia:Miscellany_for_deletion/Archived_debates/September_2019

Wikipedia contains of the mainpage, which has a large amount of traffic and from the main page the users are directed to subpages, called portals. There are portals about art, mathematics, physics and sports. Since around August 2019 there was a discussion started to delete all portals. Or to be more specific, all the 500 portals are discussed individual to keep them or delete them.

The reason for deletion is mostly the same. The page view of the portals is low, the last edit was made 5 years ago, and the interaction on the portal page is low. A similar concept to portals are wikiproject which are also subpages to coordinate the efforts about the same topic, for example about computer science. Some of the wikiprojects were deleted too, but it seems that the deletion energy is focussed first on portals.

The discussion can be described from a more abstract point of view. If the portals and wikiprojects are gone, where is the place to discuss about new articles, maintainance and peer review request? There is a place, called “request directory”. In that domain, all the domains (art, sport, science, film) are combined under a single page. That means, in the article request page, the different subjects are combined: mathematics has a subsection, literature has one and so forth. According to the pageviews, such maintenance pages are used very often by the users.

The deletion of the former portal pages is a major step in the development of Wikipedia. On the first look a portal make sense. It helps to combine different articles under a single page. The concept is comparable to a specialized library. That means, under the term “computing” only computer experts are discussing how to write content from that domain. So it's surprising that this concept has failed.

In the deletion debate, only the majority of users is pro deletion. In contrast to normal deletion debate the opposite opinion is very low. So the prediction is, that in 2020 all the other portals and perhaps the wikiprojects too gets deleted.

To understand how the meta section is working we have to take a look at the remaining Wikiproject computerscience, https://en.wikipedia.org/wiki/Wikipedia:WikiProject_Computer_science

Right now, the Wikiproject wasn't deleted, so it's a good time to observe the idea behind it. SImilar to a portal, a Wikiproject aggregates the efforts for a single domain. It described how many high quality and low quality articles are available in the domain of computer science, and it has a to do list.

The to do list contains of sections for article request, cleanup, expand existing articles, infoboxes, photo request and stubs. There is a also a list of participants who have put their username into a list as an indicator that they are motivated to contribute to the wikiproject.

The perhaps most important part of the Wikiproject computerscience is the to do list which includes article request, cleanup and so on. The interesting point is, that every wikiproject has such a list. But the list is filtered by the domain. The more general idea is, to take a look at the Wikipedia wide general to do list, which contains the same categories and combines all the domains in a single page.

So we can say, that portals and wikiprojects are obsolete and will be replaced by the general to do list to maintain all the domains like computer science, sports, films and so on.

specialized library

How can it be, that in the year 2005 many hundred of portals were established in the WIkipedia and now the same community is motivated to delete all of them? It's about how to organize knowledge outside the Wikipedia project. Suppose, there is no Internet available and a classical library is used to create an article about a subject. In the university domain, a specialized library was the normal way in doing so. The advantage of a specialized library is, that it provides only a small number of books in a single building. This reduces the costs. It's possible to create a specialized library with printed books about a single topic for example computer science. The amount of journals, books and dissertations about this subject is small. It's possible to collect all of them and put them into the library.

This was the workmode before the Internet was invented. If somebody was interested in getting an overview about the topic or likes to create a new paper, he was going to a specialized library. This was without any doubt the motivation in 2005 to establish portal pages in the Wikipedia. It copies the well working principle from the offline world.

In 2019 the situation is different. Most information are stored online, and specialized printed libraries are under pressure. What the experts users in the university are doing to today is to visit a general library and use Internet for getting access to specialized papers. The same is true for Wikipedia users. Most of them are working with fulltext search engine to get the information they need. The most used entry page is not a specialized search engine for a certain domain, but a search engine works by entering the needed keyword. The same search engine allows to browse in different subjects. This makes a domain specific portal obsolete.

Expanding WIkipedia

In the self description, a portal provides an entry page for a domain, which is adressed to the readers, and a wikiproject is part of a portal to coordinate the effort of editors to improve the articles. The idea is, that it's not possible to coordinate the maintainance of all the 5 millions articles in the Wikipedia, so the task is split into domains likes art, history, science and music.

A large scale project is Wikiproject history, https://en.wikipedia.org/wiki/Wikipedia:WikiProject_History SImilar to other other wikiproject it looks a bit inactive. But we are ignoring the low traffic and take a look what the self-understanding of the project is. The project goals are to improve the history articles in the Wikipedia by creating new ones, expanding old ones and improve the quality if needed. Also the goal is to serve as a central discussion point and to answer queries from the reference desk.

The interesting fact is, which kind of topic is offtopic at the history wikiproject. Everything outside the domain of history. That means, if somebody likes to expend an article about robotics he won't get help in the history project. This is logical but it explains what a possible bottleneck is. But let us go back to the goals. Create new articles and improve the quality of existing one is an important task in Wikipedia. It's not possible to ignore this goal but this would be equal to a failure of the Wikipedia in general. So the question is how to do this task more efficient?

The best way in doing so by formulating requests from the environment. That means, a user tries to find an article about a subject, is disappointed because the article is missing and then he formulates a request like “i need an article about topic abc. Please create one or explain to me, why the topic isn't available in WIkipedia”. There are two options how to handle such requests. One option is to focus on the subject or to focus on the request in general. A wikiproject is focus on the subject. That means, if a user formulates a request about a missing article from the subject history he has to ask the history section, and if he likes to read something about computer science, he has to go to a different wikiproject.

The more efficient way for interaction is a centralized request page, in which all the domains are combined. For doing so, the existing portals and wikiprojects have to be deleted, while the request directory should be improved and become more user friendly. A centralized request desk allows to improve all the articles in the Wikipedia.

December 30, 2019

From printed academic journals towards self-publishing authors

The term author is a new one. It is referencing to an individual who is publishing a book. The interesting fact is, that in today's academic publishing industry no authors are available, but papers are created by institutions. The so called affiliation is the most important sign to be a member of a larger institution. That means, a new paper about a computer science topic isn't created by one or two individuals but it's initiated by university department in cooperation with a journal. This is especially true for the hard science.

In the domain of human science, the situation is more relaxed. That means, in the human science individual authors are creating papers, and they doing so without asking the university first. The reason is, that in ficitional writing there is a long tradition of become an author and self-publish the work either in print format or in the internet.

Let us take a look into some recent developments in academic publishing. The upraising of academic social networks like Researchgate and Academia.edu can be called a small revolution because they put the attention on the individual author. A paper is no longer the result of group working at an institution, but it's created by individual for their own purpose. The argument against such development is very similar to authorship in general. That means, the idea that a book is created by somebody, and it's not for somebody is something which is new.

To understand the situation in detail we have to ask for the relationship between reader and author. In today's academic publishing system, the reader is the center of attention. If somebody likes to learn something about science he can visit a library, a university or a book shop. In all these cases, the reader gets a lot of help. The average library provides thousands of books, and in a university the user can decide between many hundred of lectures. Now we have to focus on the other social role. Suppose, the idea is, that somebody likes to write and publish a book. Which options are available? The amount is small, or to be more specific, there is no way, that somebody is allowed to publish anything, especially not if the topic is within science or technology. The reason is, that publishing in the ivory tower works not by individuals, but with larger groups for example it's driven by a company, or by a journal.

It's not very hard to predict, that in the future the situation will flip. The individual author will become the top priority and the reader gets ignored. This has to do with the increasing amount of authors. That means, if all the people are authors, all the people likes to publishing something on an individual basis. The boom of the self-publishing market is only the first step in that direction. Self-publishing in the new, means usually to publish ficitional books, but not knowledge about science and technology. The reason is, that in the fiction segment it's very common that the authors stays in the focus who.

The open question is, how authorship will become dominant in the non-fiction segment. One way in doing so is to define rules under which publication of non-fiction / scientific books make sense and under which situation not. A possible rule is, that only the publication of detailed specialized knowledge make sense. A non-fiction writer is asked to write a book which fits into a department library of a university, but not into a universal library. The idea is, that a document about the dewey classification 500 is useless, because 500 isn't specific enough. It's reserved for general information about natural science and mathematics. The better idea is to write a book about DDC 519.233 which is about markov processes as a subsection of probability theory. The value of non-fiction information depends on how deep in the DDC22 tree the book is located.

December 28, 2019

Analyzing the bottleneck in academic publishing

To predict the future of academic publishing it make sense to focus on conflicts in the existing library system. In general the situation can be divided into the time before the advent of the internet and the time after the 1990s which is working with electronic communication.

Academic publishing before the internet was located around special libraries. The reason was, that only special library are able to provide detail knowledge. In a single building, all the books and journals about a subject can be collected. With a general library this was not possible, because the amount of shelfs and the costs would explode. If no internet is available and printed books are the only medium, a special library fulfills the needs of professional engineers well. The user comes with a concrete specialized problem, and the library provides the answer to the question.

During the transition from printed journals to online journals there was a bottleneck visible. Most specialized libraries have failed to transfer the concept into the internet. Some library portals are available in the internet, which are providing access to a low amount of ressources from a special domain. These website have failed. The reason is, that internet users are not interested in using 100 different websites, but they are interested in a meta-search engine which provides as much information as possible. As a result, not specialized libraries but information aggregation search engines have become famous in the internet age. The most famous one are Google Scholar, and Elsevier based universal search engines.

It seems, that in the internet age, the concept of a specialized library has become obsolete. This has produced a lot of stress to the system because there is a need for something which is new, but nobody knows what the future library will look like. What is sure was the old outdated concept of specialized libraries. The main idea was to work with constraints which are reducing complexity. In a specialized library about computer science, only books and journals from this subject are welcome. Everything else gets ignored and is called not relevant. That means, if a user in a specialized computer science library likes to read an article about music theory, the request was denied because it doesn't fit to the core subject. Addtionally, all the users in a specialized library have an expert background by default. They are experts for a certain subject but not informated about all the other knowledge in the world.

The main benefit for complexity reduction is to lower the costs. If books outside a specialized domain are ignored, they do not have to stored in the shelf. This saves money, human labor and physical space. The combination of low costs plus indepth information access was the success factor for special library.

The working hypothesis is, that with the advent of the internet, the restrictions of the past do not make sense anymore. The problem is, that the complexity has exploded. The world of a special library was easy to understand. But if all specialized libraries are merging together into a single universal library, it's unclear what the system is about.

Discipline-oriented digital libraries

Since the 1990s, some attempts were made to build so called “Discipline-oriented digital libraries”. The idea is to transfer the concept of a special library into the internet age. This idea make sense but fails at the same time. From the perspective of a printed special library it make sense. A special library is the core of academic infrastructure. It's major bottleneck is, that no fulltext search is possible. Providing the same information in the online format makes a lot of sense for the users.

At the same time, these websites have failed. Because the result is, that lots of different search engines are available. What the user likes is a meta-search engine which can search through all the content. One option in doing so are information aggregation websites from Elsevier and Springer which combining the content from different special libraries. The disadvantage is, that a publishing house is different from a library. So the question remains open, how an internet version of a special library will look like.

The reason why it make sense to analyze classical special libraries in detail is because they can answer the question what academic publishing is. Academic excellence and a special library is the same. So what exactly is the difference between a normal library and a special library? It has to do with a certain sort of information. In a normal / general library, the books are about fictional topics. For example the novel Anna Karenina from Leo Tolstoy is an example. A specialized library won't collect such a book, even if it's world literature. A second property of a special library is, that it provides in depth knowledge about a concrete domain.

So we can summarize that specialized library are focussed on non-fictional information with detailed topic in mind. If a document contains lots of special vocabulary spoken by experts on the field, and provides non-fictional content, then the document is scientific.

Building a digital discipline oriented library

Reproducing the specialized focus of an academic library in the internet age can be realized with tags. The difference between a fictional book from a general library, and a scientific book is, that the second one is tagged differently. A non scientific book can be tagged with “novel, adventure”, while an academic book is tagged with “non-fiction, computer science, neural networks”. What academic users are demanding, is content which is tagged in a certain way.

Specialized libraries and special journals are creating the reputation by focus on a certain topic. Academic work means, to ignore most of the information and oriented only on a sub part of the problem. Roughly spoken, a journal which is publishing papers from different disciplines is not an academic journal. Only a journal which is dedicated to a concrete issue is able to create a reputation.

At youtube there is a video available which introduces a science library from the late 1970s, Life Sciences Library 1979 - Restored Version, https://www.youtube.com/watch?v=xe0lFUlz-io Even if the library is not very large and it working with printed index cards, it's a state of the art library even for today's needs. That means, if the internet is offline, a specialized library would provide the same or even better information available online. The low amount of books and the missing fulltext search is not a problem in a special library, because the content can be explored manually.

science library are the core of academic excellence

In the public perception, special libraries are often ignored. They doesn't have prestigious buildings and in contrast to universal libraries the amount of books they have to offer is small. Most special libraries are small in the size. That means, the amount of employees is below 100, and it's located in a single building. On the other hand, special libraries are more important for researchers than large universal libraries. What researchers are doing at foremost is to focus on a concrete topic. This is equal to become an expert. And special libraries are needed by the experts for reading new information.

A special library is the same like academic working. It's not possible to become an expert for everything. But there are domains like physics, mathematics, biology and so on. It make sense to collect all the information in a single library and prevent of creating larger universal spaces. This is not located in a certain mentality, but the simple reason has to do with costs. A special library produces lower costs than a universal library. At the same time a special library creates a high entry barrier. Users from a different department are not allowed to enter a special library. Only mathematicans can visit a mathematical library.

In the electronic age the borders have blurred. The idea is to get access information worldwide and from all subjects under a single search engine. This new understanding of a knowledge hub is working different from the former science library. And the question is how to maintain the normal specialized quality standard in the electronic age.

Let us give a concrete example. The hot topic in the age of Open Access are so called predatory publishers. That are electronic journals which have lowered the entry barrier. Predatory journals are perceived by academic as low quality journals. But what if someone creates a specialized predatory journal? That is a low entry barrier journal which is dedicated to a special topic and accepts only manuscripts from this domain. Such a predatory journal has to be called a normal academic journal, because it provides this sort of information which is needed by experts. Or let me explain it the other way around. If a researcher gets specialized information about his domain which are uptodate, the researcher is happy. No matter if the journal is called predatory or not.

In contrast, a fake journal is equal to the attempt to describe a subject in a general way. That means, by avoiding the specialized vocabulary of the expert. Mainstream information which are targeted to a larger audience can be called a fake science journal because they do not fulfill the needs of academics.

Making Wikipedia more accessible to the public

Introducing Wikipedia in the year 2019 to a wider audience is no longer needed. The website has reached the top10 of the Alexa statistics and can be called a success. Similar to encyclopedic projects in the 18th century, Wikipedia has adopted the latest technology and tries to enlighten the society.

Around the Wikipedia project there is a major concern, which is located not in the content itself, but in the documentation what Wikipedia is about and how to contribute to the project. The problem is, that many help pages are created and additionally in the academic literature a large amount of papers about the Wikipedia were written. It seems, that even for bibliographic experts, it's hard to explain what Wikipedia is and in which direction the project will evolve in the future.

The literature about Wikipedia can be divided into two groups. The beginning was dominated by papers about the project itself, and it's comparison to existing encyclopedia projects like Britannica. SInce 2010 a new sort of papers was written which is focussed on the conflicts in the project. Introducing this sort of text needs a detailed explanation.

In general, there are two options available to describe a project. The first idea is to assume, that Wikipedia is a blackbox which can be communicated to the outside world. The second idea is, to describe Wikipedia as a living system which is powered by internal conflicts. A typical example for a conflict is an edit war, banning a user from the project or deleting an article. The interesting fact is, that smaller wikis which are created by a single admin, doesn't have these sorts of conflicts. In a single user Wiki, the only possible conflict is available between the human user, and the mediawiki installation. For example, the user tries to format a heading in bold, but the syntax parser produces an error.

In the Wikipedia project are more complicated sort of conflicts is visible, which has to do with interaction users which are trying to achieve different goals. The most obvious conflict is between a user who likes to enter new information into the Wiki and the admin, who likes to prevent this, because he classifies the edit as vandalism. The amount of conflicts in the WIkipedia isn't researched very well. In the early literature, the assumption was, that conflicts can be ignored. From a game theoretic perspective, it make sense to monitor the conflicts in detail, because they are the explaining what the rules in the system are.

It's important to know, that Wikipedia is working different then it was described in the help section. That means, the rules of the Wikipedia game are not described explicit but they are communicated in the conflicts which break out and which are solved in a certain fashion. Describing these conflicts allows to identify the shared goals of the users and under which cases a stress from the outside is available.

How exactly are conflicts solved in the Wikipedia? The answer is, that the users in the system are anticipating the reaction of the community and this allows them to adapt their individual behavior. That means, a conflict is located on the time scale and it produces reactions in the future. In most cases, the prediction of future behavior is based on looking backwards in the past. Longterm admin users have a large knowledge what the workflow was for certain article conflicts. This pattern is used to generate a certain behavior in the now. A behavior consists of entering text and pressing delete buttons.

Creating a Wikipedia article

... from scratch doesn't make much sense. Because the user has no idea, which topic is important nor how to format the paragraph so that Wikipedia is happy too. The more elaborated way in creating content for Wikipedia is to search for text already written on the own home directory. In the best case, it's draft for an upcoming wikipedia article, written 2 months ago, but never uploaded to the encyclopedia. Such a draft version can be extended with two literature references, and after proofreading it's ready for the sandbox. This tool allows to check if the paragraph is well formatted, and then it can be copied into the article space.

The newbie would assume, that after uploading new content to the Wikipedia, an edit war will start in which a powerful admin collective will go through every referenced source and will ask if the user is already familiar with the topic. Such a scenario is available for mainstream articles which have a high pageview statistics, but the edit in normal scientific article in the encyclopedia is mostly ignored. That means, the content is uploaded and nothing will happen. If the edit looks not as maximum spam, but seems to look halfway informed, the edit will be accepted as valid. The reason is, that most articles in the Wikipedia doesn't have any contributors at all. That means, the last substantial edit was made 2 years ago, and if someone likes to add a small paragraph Wikipedia won't reject it.

Sure, it's important to write accurate sentences and provide quality ressources in the footnote section, but in general the quality standard is only on the average level.

Academic publishing isn't invented yet

If someone tries to identify the best practice method in academic publishing he will recognize that the subject is highly controversial. The reason is not, that each journal and each university has a different opinion but the major problem is, that so called academic publishing is invisible. Invisible means, that there is no history of publishing available which can be described and reproduced. The missing data from the past can be observed if the annual papers who get published are counted https://www.scimagojr.com/countryrank.php

In the year 1996, all the countries in the world have created less than 1 mio new papers. The United states holds the record with 350k and the other countries have each under 100k. Nearly all the papers in the 1996 were published in the printed form by larger publishing houses. This statistics doesn't describes a culture of publishing something, but a culture of doing the opposite. Let us make a thought experiment. Suppose we are take a look back into the year 1996 and try to read one of the published academic papers. How can we do so? One option would be to visit a university library. If we are going to a library in Europe, in China or in southamerica, the chance is high, that even the largest library in a country has only subscribed to local journals written in the local language which is not English. That means, in a Italian university of the year 1996 there was no Internet access, but all what the reader can expect are some journals in the archive written in Italian. Even if the user is asking for fulltext access to advanced research he won't be able to read such papers. It's also not possible to create new content.

According to the URL, the total academic output of Italy in the year 1996 over all disciplines was only 40k papers for all subjects. That means, if the user is interested in reading the latest research of this topic, he will get from the Liberian a list of 2 journals and with a bit luck, these journals are available in the library. They will fit on single desk and it can be read in one afternoon. The problem is not to describe the workflow of reading and writing academic information in the year 1996, but the problem is, that in this time no such thing like academic content was available. That means, there was an absence of excellence.

In the year 2018 the situation has increased a lot. Today, there is the internet. If we are repeating the thought experiment and travel to a library, the user gets access to online repositories from worldwide publishers. He is able to read worldwide information which are provided as paywalled and open content as well. In the year 2018 the worldwide paper production was around 3 million documents, which is much higher than in 1996 but it can't called academic publishing, because very similar to the past, most information are not written yet. Publishing a paper has a low priority. It is something done by publishing house and no standard workflow is available. What most professors are doing is not to create and upload new information but they are doing nothing. Most academics didn't have written a paper in the last year, and if they have done so it wasn't uploaded to the internet because of different reasons.

The reason why it's hard to define academic publishing is because it wasn't invented yet. What we have seen in the year 1996 was publishing with printed journals with a low amount of documents. And what is available in the year 2018 is a complete different publishing mode which can't be extrapolated into the future. The only thing what is sure is, that in ten years from today, academic publishing will work different from today. Inventions like fulltext search, a reputation management or peer review wasn't invented yet. That means, the subject can't be descibed by looking backward but it has to be invented for future needs. Some ideas who future academic publishing will look like are available today. A combination of a preprint server, fulltext search engine, reputation management and grounding in physical locations like universities make sense. The open problem is how to combine all these stakeholders and make the system open as possible.

One option in approaching the topic from a conservative standpoint is to ignore future needs and claim, that the publication system of the year 1996 was working great. Under this assumption it make sense to focus on a few printed journals which are publishing less than 1 million papers per year with a well working quality control. The future academic publishing system will look the same like in the 1990s before the internet was invented. Which means, that no content at all is created and the world has no access to the information. The question is, if such an idea make sense for the internet generation as well? The main advantage of the 1990s publishing model is, that it helps to reduce complexity. Instead of trying to improve something and invent lots of new workflows, everything remains stable.

The assumption is, that the domain of science is walking slowly forward and there is no need to increase the paper count. Important papers are written in the printed format and they are located physically to a library. The total amount of researchers in a country is small, and if they want to explain something to the public, they can invite the local newspaper into the researchlab which can write about something discovered recently. The newspaper acts as a buffer between the scientists and the public and what most of scientists are doing is to proof that everything is at the right place.

The unanswered question is, what is the role of science and technology in the world? Does a modern society needs a progress at all? Is there a need for increasing the amount of scientists and does it make sense to publish 3 mio paper a year?

During the decades the role of a library has changed. The importance has grown and today, the world has a higher demand for academic information than in the past. 50 years ago, the term library was referencing to book shelf which consists of 40-100 books mostly with fictional content, novels and poems. Today, a library is equal to a scientific library and sometimes, it's equal to an online library which provides fulltext access to the latest research papers. The definition what a library is, has become more quality oriented. Today, a bookshelf with 40 fictional books can't be called a library, but it's a joke. The reason is, that the value of these books is low, and doesn't fit to basic needs. 50 years ago, such a bookshelf was equal to a library used by educated people. The reason was, that in the past it was rare to have access to books at all. And reading fictional books is better than reading no information at all.

Libraries

Knowledge production is working as an asymmetric system: one side accumulates all the wisdom, why the other side doesn't have access. Before the internet age, libraries were the hub of academic knowledge. They are providing books to a small amount of people. Reading the books makes the people educated.

There are two sorts of libraries available: general libraries which are accumulating as much information as possible. They are collecting different languages, fictional and non-fictional books from all sort of topics. Building such libraries in the 1990s was very expensive, because lots of storage space was needed. The more interesting sort of libraries are specialized libraries which are dedicated to a single topic and collect indepth information about a subject. Special libraries are the perfect choice for academic purposes. They are able to overcome the limitation of the printed book format.

A typical example is a music library. Even if all the information is based on printed information such a lbrary would provide in the 1990s an indepth knowledge about the topic. Or let me explain the situation from the other point of view. General libraries in the 1990s were available but they are ignored by academics. Even if a general library has a large amount of ressources, it failed to provide indepth information about a subject. That means, the needs of researchers doesn't fit to what a general library has to offer.

special libraries

The most surprising information about special libraries is, that even before the internet age, they were perceived as powerful institutions. They are working similar to normal libraries with printed material. That means, books and journals are stored in bookshelfs, and if somebody likes to read it, he has to visit the library. What makes special libraries useful is, that in a single building all the information is stored. The researcher has a concrete question about a specialized subject and the library will help him.

That means, it was possible to do advanced research and find out something new without using the Internet. Even today, special libraries are very important for academics, because they have all the information in a single place and they are providing fulltext information. A nice thought experiment is to digitize a special libraries. Such a project has not the attempt to convert millions of books into digital information but only a few from a limited domain.