June 06, 2018

Which programming language should i learn?


The question itself is a bit outdated. It was relevant in a time in which sourcecode was not available public so the only way to get new software was to program it by own. Nowadays the typical user isn't programming anymore instead he selects an app from the Android/ios app store, he buys a box in a computerstore or he installs new open source software with the dnf package manager in Fedora.
But sometimes there is a need for programming software from scratch, in most cases for educational purposes. And here has the user the problem that it is unclear which programming language is the best. The most important information is, that between the different types of language the differences are very small. That means all the language are based on textfiles which are stored in the memory and the files containing commands, classes and variables. A second important information is, that in reality nobody really programs in a certain programming language. What the Stackoverflow and github users are doing is to program with a certain library, purpose and operating system.
For example, the goal is to program a pacman clone. The question is not, if Javascript, C++ or C# is the best choice. Instead the programmer will ask if the software should run in Android, on a desktop PC. And he will ask if the graphic contains only Ascii or 2d or 3D. That means, the choice for a certain programming language is mostly a random choice and has to do with software already there.
Let us make a more simpler example. Suppose the goal is to program a prime number generator. The interesting fact is, that in every language the result will look equal. In most cases, the correct answer will look like the sourcecode in rosettacode.org/ and there is no real difference between Python, C++ and Java. The only choice available is between procedural and object-oriented programming. And indeed, this is a question which has to be mastered individual. That means, a procedural primenumber generator looks different from his OOP-counterpart.
At the end, I want to explain who a programmer is able to master a language. Instead of spoken language like English, it is not possible to speak a programming language fluently in a sense, that somebody can learn the language and use it for real problems. The workflow in reality works differently. It is grouped around the edit-compile-run cycle. That means, the programming is taken a codesnippet from the last week, for example his “hello world” example in C++ and is trying to extend the code into a prime number generator. On the way, he recognizes that he gets problem with the int-datatype. Now he has a clean problem which can be solved by asking an expert or asking google. After he has mastered the subproblem he writes down the sourcecode.
To make the point more precise. The programmer is not able to solve the problem before he gets started, instead he is trying out always things he doesn't know. That means, he is interested in his personal inability to handle a programming language or a mathematical problem, because this gives him the opportunity to interact with a higher instance, for example with a programming expert, with Stackoverflow or with the Google Search box. That is the reason, why programming is often called a social activity, because the programmer imagine in the workflow what a certain person will say about his code. For example, the student imagines what the professor will say, if he see his prime-number-generator in C++. Is he fascinated, will he understand it?
In reality, programming is not grouped around languages, but around programming groups. It is bit like a role-playing game. The student is part of a certain group, or he hopes to become part of it, and now is trying to anticipate the group behavior to get a higher position.

The github microsoft deal makes sense


According to the latest news headline, Microsoft has acquired github for 7.5 billion US$. Would this result into a monopol in software development? No, because the monopol was there before the merger. A short look in the latest stackoverflow survey shows us, that 80% of all developers are programming exclusively under MS-Windows. That means, they have Visual C# installed and programming in the IDE C# apps, Python programs and Java apps. Only a minority of the programmer is using Linux or Apple operating system. On the github platform the situation is equal. Nearly 100% of all C# projects are written for the MS-Windows operating system, 99% of the C++ games, and perhaps most of the HTML5 and Python projects were developed in Windows too. So, if Microsoft comes to the conclusion to buy the portal it makes sense for the users.
In reality, Linux is not very common in the programming community. Worldwide around 5 million Fedora installation are out there. That means, core- Open Source developer which are using an Open Source operaring system and want to program an open source project are a minority. Github and Stackoverflow were never independent or Linux centric, from the beginning it was a Microsoft play.
Sure, some github repositories were created by Linux developers. That means, a MS-Windows operating system is not needed to compile, install and improve the sourcecode. And many repositories have an Android background. But, average repository has a Microsoft background. That means, the developer is sitting behind a Windows 10 machine and he programs in Java or C# which are both proprietary languages. I would guess, that most users have a problem with recognizing the reality. They want to believe, that they are hardcore Linux users, only because they have installed the Linux subsystem for Windows, but in reality they have absolutely no idea about Open Source software or the Red Hat ecosystem.
I think, it was a great move from Microsoft to acquire the number one Open Source developer platform, because this helps to compete with Red Hat. Otherwise, Microsoft would loose much more market share to the inventor of Linux.

Multi languages in European science


It was a pleasure to read the article. The writing style was very fluent, no grammatical error were in there and the idea itself is interesting. I've also recognized that former local languages like German, French, Italian and Mandarin are very rare in the scientific community. So it would be a nice idea to promote the usage of these language. That means, if somebody has written a paper, for example in German about a current topic in science, he should become top-priority compared to a normal submission which was written only in English. I think, the local European languages should become a much broader role, because otherwise they will die. It is important to train at universities with the student to write in their local mother tongue. It must not always be German, the Chinese language is also an interesting example. Writing a paper in Chinese can be a demanding task, especially for people who are not familiar with it. If somebody is able and motivated in doing so, he should be supported by the community, which means by his university.
Another important point is, that for many English words no translation is available. The idea of French to translate computer into ordinateur is a good example, how a modern scientific french can be look like. That idea should be practiced in other languages too. This helps to enrich the vocabulary, and to make a local language more attractive for the students. A starting point is to translate the term “machine learning” into non-english languages, for example into “apprentissage automatique”, “maschinelles Lernen” or 机器学习.

June 04, 2018

Bob and Alice are building Nanorobots with a special power supply


Alice: [is using a modem dialup connection]
Bob: [has a preinstalled UNIX operating system]
Alice: Connection established ...
Bob: Hi Alice, what's going on?
Alice: Everything is fine.
Bob: ...
Alice: I only wanna say hello.
Bob: No you don't. I've read your message in the bulletin board system ...
Alice: You mean the message with the biomedical announcement?
Bob: No, the posting with the basketball results ;-) Hey come on ...
Alice: ... ok my problem is simple. The nanobots are not working. And i have absolutely no idea why.
Bob: Now we going into the right direction.
Alice hm.
Bob: You're nanobots have problems with the energy, right?
Alice: That's true. The onboard generator isn't working with maximum productivity.
Bob: I've read the specs, it is using radioactivity?
Alice: What else?
Bob: I only want to play it safe ...
Alice: Something is wrong with the decay rate.
Bob: Do you have asked the Doctor for help?
Alice: No I don't trust him anymore.
Bob: ???
Alice: It is possible that he is no longer on our side.
Bob: No, I'm sure we can trust him.
Alice: Why?
Bob: Because he has invented the docking maneuver in Nanotechnology.
Alice: yeah, but his knowledge about “liquid drive” is low. And he is working for the other side.
Bob: You mean, he is part of a much broader conspiracy?
Alice: I'm talking about non-proofen failures in the last year. For example the energy loss under 20 degree temperature.
Bob: That would explain something. Perhaps we can agree that we need further investigations?
Alice: Sure.
Bob: I've got on the other line a critical request.
Alice: Ok, let us continue the chat in the next week.
Bob: You're welcome. Bye.
Alice: Bye.

Realistic estimation of thesis writing productivity


A good average is 4 hours per page, which is needed in phd / thesis writing. On the internet, sometime higher values are discussed. For example i've found many posting in which the total amount for a phd thesis was given with 3000 hours over 7 years. And the result was a 200 pages long phd-dissertation. According to a short calculation this is equal to 15 hours per page.
But how can we minimize the value? At first it is important to know, that in reality the effective productivity is in between these borders. That means, it is not less then 4 hours per page and not more then 15 hours per page. The main problem with thesis writing is, that the writing itself can only be done as a result of studying. That means the workflow consists of reading existing literature, thinking about, making experiments and writing down new information. Sure, it depends a bit on the needed quality and also the subject, but there are some activities which are equal in all domain. In the minimalistic form the workflow contains two parts:
1. reading existing literature
2. writing the own text
This is called minimalistic because all the other steps like proofreading, formatting, doing experiments and filtering out not wanted results are missing. It is not possible to shrink the workflow further. So I'm sceptical if it is possible to need less then the above mentioned 4 hours per page. The only exception is Artificial Intelligence, for example the IBM Watson software which is – in theory – able to generate short text. Such a textgenerator needs indeed less then a second for producing content, but in reality most papers are written by real humans.
The problem is, that even with modern technologies the productivity is not very high. And even experts in academic writing are not able to increase their productivity further. That means, even if somebody familiar with a subject and an expert for English language, he needs some time to write a text from scratch.
But in reality, the bottleneck isn't a real problem, because it is possible to increase the number of students. If one student can write in one week a 5 page long paper,then 10 students can write in one week 10x5 pages and so on. I'm pessimistic if it is possible to increase the subjective productivity of an author, but it is always possible to increase the group-productivity. The average author is able to type in an average quality paper. It is not a mark 1 but it describes a topic. If there is a need for more academic paper this can be fulfilled by more students who are writing such content. The interesting fact is, that the world has around 7 billion people which is in theory an unlimited reservoir of academic writers.
The funny thing with any electronic document is, that after it is created once, it can be copied many times. A pdf file doesn't loose his quality if it is downloaded 1 million times, it remains the same data. If the content is available at the internet, all the people can profit from it. That means, even if the generated content has a low quality, is about the wrong topic or was written in the wrong language it will always improve the Gutenberg-galaxis. It is not possible that somebody can weaken the overall accumulated knowledge by adding new information. The only possibility is to reduce the amount of information, but that is technically not possible.

Some insights in the world programmer population


According to the Public Relation newsarticle nearly all countries in the world are heavily interested to educate the people in programming skills. China alone has announced massive investment, but also the US and Europe have millions or billions of people with programming skills. Unfortunately, this description is some kind of cyber-myth which has absolutely no relevance for reality. The good news is, that nowadays it is relatively easy to get valide information about how many programmers are available worldwide. The major and only programming online forum available is stackoverflow and on their statistics section they have a detailed number about the users, https://stackexchange.com/leagues/1/year/stackoverflow Right now, they have 200k active users worldwide, that is the number of user accounts who has posted at least 3-4 postings in the last year.
That's it. The number are not billion, it is not million, it is 200000 worldwide. We can now discuss about the detail, perhaps in which country the programmer life or if they are familiar with Java, C or something else. But the fact is, that there is no hidden programmer army above the 200k level. The second fact is, that nearly everybody who is starting programming is listed in the statistics. That means, only a subpart of the 200k people have ever written real big programs, most of them are amateurs and hobby programmers who have only posted a novice question and increased their reputation counter.
Now let us take a look into mainstream public relations myth about software engineers. According to https://en.wikipedia.org/wiki/Software_engineering_demographics it is estimated that many million people are able to program:
“As of 2016, it is estimated that there are 21 million professional software developers.”
I would say, we have two numbers: Stackoverflow says that the number is 200k, WIkipedia talks about 21 million. I would guess, Wikipedia is wrong. Because all the million hidden software engineers need two things to be valid:
1. they need to program some kind of software
2. they need some kind of online forum to discuss problem
Because there is no hidden software pool and also there is no hidden online forums in which millions of programmers can discuss problems which are unknown for Stackoverflow, we must conclude, that only the Stackoverflow number is correct. That means, if somebody is interested seriously in software development he will post questions online, otherwise he will do nothing.
A nice side-note is, that even big companies in the gaming-industry who are programming games by hand have only a small number of employees. If a software-company has more then 100 employees it can be called huge. That means, in the 200k active Stackoverflow users, all the professional guys from Microsoft, Electronic Arts and Apple are already in-there.

May 31, 2018

How long does it take to write a phd thesis?


In the domain of computerprogramming there is a benchmark available which measures the productivity of an average author. No matter how good somebody is, or which programming language he is using, he will reach only 10 lines of code per day. In the domain of thesis writing there is a similar benchmark available. Here is the question how many hours somebody needs to write a single page in his phd-thesis. The average value is 4 hours per page. That means, if the aim is to create a 150 page long thesis, the average author needs 600 hours.
The interesting aspect is, that we can discuss about this benchmark. Perhaps if it is to demanding and in reality the author will need 8 hours per page. But according to different discussions in the Internet, the number of 4 hours per page is fair to describe the current situation. It predicts quite good, what is possible and what not. Let us make a thought experiment. Somebody needs a 10 page long article for a journal submission. According to our above cited formula he will need 10 pages * 4 hours = 40 hours to write the paper. Is it possible to reduce the effort dramatically? No, like in the example with the programming productivity, it is a fixed value which is surprisingly constant over periods. What is possible is to try to increase the value, for example with a better wordprocessing software which allows faster formatting of a paper. But at the end, the bottleneck is the author himself. His ability to read the given literature, doing experiments and understand a domain is limited.
To make the situation a bit more realistic we can search for any dissertation on the internet. Perhaps the pdf document contains of 150 pages. We can download and read the dissertation in under 5 minutes, but for creating it the author had invested lots of energy. In most cases, a phd thesis was written over 2 years. It is the result of an ongoing research effort, which takes not only weeks but months until it was finished. And this is perhaps the most important reason for the Open Access movement. If a single phd thesis needs many months to write, it is a waste of energy to hide it from the public. The best time-saving advice is to not write any academic paper. That is only possible if paper written in the past are available in the internet, so there is no need to write the same information again.

How much researchers are enough to bring AI forward?


The ideology behind Open Science is to transform the masses into a scientific workforce and motivate millions of people to work together. They are distributing content like the dancers on the Love parade in Berlin which has over 1 million people on the same time who are collaborating. But, do we need such amount, or wouldn't it enough to mobilize only a small amount of people?
The famous Artificial Intelligence forum https://ai.stackexchange.com/ has only a limited amount of regularly users. According to the last statistics, in the last year only 50 people have posted some content to the forum. Surprisingly this small amount was enough. The forum contains lots of answers from different subjects, and incoming new questions gets an answer reasonable fast. The hypothesis is, that 50 serious users who are contributing to a forum or an academic journal are enough to bring a discipline forward.
Suppose we have 50 users, who are writing academic papers on the subject of Artificial Intelligence. Each user is able to write one paper a month. After one year, the group has produced 600 papers. It is not a second Arxiv.org repository but it is enough to bring the subject forward. So perhaps the ideal Open Science community consists not of million scientists, but only of 50? The question is not how to attract the whole planet to attend an online forum, the question is what the upper limit is. I would guess, that https://ai.stackexchange.com/ can handle perhaps 200 users who are posting questions and answers there. More wouldn't improve the situation but result into chaos. Today's 50 users are a bit small, but it is not advisable to increase the number to 500 or even more.
I would guess, that in other disciplines like economy, literature, biology and medicine is the situation equal. A single researcher is not able to run a forum and a journal, but 50 people who are working together are more then enough, and increasing the number of scientists to the range of 1000 or more would not help to improve the situation.