Showing posts with label Natural language. Show all posts
Showing posts with label Natural language. Show all posts

August 08, 2026

Sprachspiele im Kontext von Robotik

Die Erforschung künstlicher Intelligenz drehte sich lange Zeit um die Frage welche Art von Softwareframework, Algorithmus oder Programmiersprache benötigt wird um denkende Maschinen zu realisieren. Die Annahme hinter dem Cyc Projekt von Douglas Lenat lautete, dass eine Ontologie der Grundbaustein sei, das Cam-Brain Projekt von Hugo de Garis unterstellte dass im Kern ein neuronales Netz benötigt wird während Edward Feigenbaum vermutete dass sich Künstliche Intelligenz mit Hilfe von Expertensystemen realisieren läst.

Im Laufe der Zeit wurden sehr viele gegensätzliche Technologien und Annahmen entwickelt um Künstliche Intelligenz zu verwirklichen. Ein Sprachspiel im Sinne von Ludwig Wittgenstein ist nur ein weiterer Vorschlag unter vielen. Dennoch lohnt es sich, das Thema nähter zu untersuchen. Weil das Prinzip von Sprachspielen deutlich anders ist als z.b. ein Expertensystem oder ein neuronales Netz.

Sprachspiele sind ein formalisiertes Interface zwischen interner und externer Realität von Systemen. Es ist ein Test, ob die Spielteilnehmer Sprache erzeugen und verstehen mit derer man die externe Realität beschreibt. Ein typisches Beispiel ist das "Ich sehe was du nicht siehst Spiel". Dabei beschreibt Person A ein Objekt aus der Realität anhand von Eigenschaften, er formuliert aussagen wie "das Objekt ist rund, steht in der Küche, hat Beine" und daraus folgert Person B "es ist ein Tisch".

Die Gemeinsamkeit von allen Sprachspielen ist, dass Sprache in einem interkationsspiel genutzt wird um auf die Wirklichkeit zu referenzieren. Es geht immer darum, Dinge abzufragen die in der Realität vorkommen, Anweisungen geben was in der Realität zu tun ist oder sonstwie Sprache und Wirklichkeit in Beziehung zu setzen. Sprachspiele sind eine gute Möglichkeit eine Fremdsprache zu erlernen und dienen dazu den Wissensstand zu erfassen. Wenn eine Person oder ein Computer in einem Sprachspiel eine hohe Punktzahl erreicht, ist diese Person mit der jeweiligen Sprache vertraut.

Anders als die eingangs erwähnten Technologie zur Realisierung von künstlicher Intelligenz wie Onotologien oder Expertensysteme sind Sprachspiele nicht technisch definiert sondern haben ihren Ursprung in der Philosophie und der Linguistik. Es geht um Themen wie interaktivität, Sprache, Wirklichkeit. Die Annahme lautet dass innerhalb dieser Begriffe Künstliche Intelligenz möglich ist, das also denkende Maschinen ein Interface zwischen internem System und externer Realität sind.

Seit den 1980er wurden Brettspiele als Testumgebung für denkende Maschinen genutzt. Schach ist ein sehr altes Beispiel für das Künstliche Intelligenz realisiert wurde, aber auch andere Spiele wie Tic Tac Toe, Dame, Backgammon und Go sind klassische Umgebung zur Erforschung von KI Algorithmen. Leider haben diese Brettspiele den Nachteil dass sie nicht gut nach oben skalieren. Ein Computer der perfekt Schach spielt ist nicht automatisch im Stande einen Roboter zu steuern. Deshalb eignen sich die erwähnten Brettspiele nur sehr eingeschränkt dazu KI näher zu erforschen.

Sprachspiele kann man als neuartiges Gesellschaftsspiel verstehen was ähnlich wie Schach Regeln folgt aber viel besser nach oben skaliert. Ein Computer der das "Guess what" Sprachspiel beherscht ist zugleich auch in der Lage mit diesem Wissen einen Roboter zu steuern. Scheinbar können Sprachspiele die Kernidee von Künstlicher Intelligenz viel besser formalisieren als frühere Brettspiele.

July 24, 2026

Die späte Entdeckung der Schrift im Kontext von Robotik

 Die Schrift ist eine sehr alte Erfindung der Menschheit. Die erste Bilderschrift, die Ägyptischen Hieroglyphen entstanden um 3200 v. Chr. Insgesamt ist Schrift und natürliche Sprache sehr detailiert erforscht. Es gibt umfassende Wörterbücher, Darstellungen welche die Geschichte der Sprache zeigen und Untersuchungen bezüglich Wortherkunft.

Vereinfacht gesagt sind Wörter Referenzsysteme zur Realität. Substantive stehen für Objekte wie "Himmel, Tisch, Apfel", Adjektive stehen für Tätigkeiten wie "Laufen, springen, geben" und Adjektive werden als Eigenschaftswörter verwendet wie "gelb, groß, schnell, feucht". Das Wissen bezüglich Wortarten und die Nennung von Beispielwörtern ist banal, allerdings nur für die Sprachwissenschaft selber. Im Bereich Computerwissenschaft und Mathematik wurde natürliche Sprache lange Zeit ignoriert. Es gab zwar Versuche chatbots zu programmieren, aber das war nur ein Teilbereich der Künstlichen Intelligenz.

Erst in jüngerer Zeit stellte sich heraus, dass natürliche Sprache womöglich das fehlende Puzzleteil darstellt mit der man künstliche Intelligenz inbesamt realisieren kann. Und zwar indem man Sprache als Technologie verwendet. Insbesondere dessen Eigenschaft auf die Realität zu verweisen macht es zum idealen Abstraktionsmechanismus. Es müssen keine neuen Sprachen erfunden werden sondern vorhandene Sprachen wie English, Deutsch usw. bieten bereits ein umfassendes Vokabular was von Robotern ähnlich wie Menschen verwendet werden kann. Alles was eine Maschine dafür benötigt ist eine Übersetzungstabelle von Bildern zu Sprache und in umgekehrter Richtung. Mit Hilfe dieser Bild zu Wort Tabelle kann man einem Roboter ein Kommando geben wie "Fahre zum Tisch". Und der Roboter übersetzt den Satz dann in Bilder und in Aktionen.

Bis ungefähr zum Jahr 2000 hat die KI Forschung nach Algorithmen gesucht mit deren Hilfe sich denkende Maschinen konstruieren lassen. Typische Algorithmen waren Lernverfahren für neuronale Netze, SLAM Algorithmen zur Selbstlokalisierung, A* Pfadplanungsalgorithmen oder Momdel predictive control Algorithmen. Die Annahme lautete jeweils dass mit diesen Algorithmen denkende Maschinen konstruiert werden könnten. Diese Annahme ist jedoch falsch. Es liegt zusätzlich der Verdacht nahe, dass es generell keine Algorithmen gibt, die Künstliche Intelligenz erzeugen, weil ein Algorithmus per se nicht mächtig genug ist um Roboter zu steuern. Was man stattdessen verwenden könnte wäre natürliche Sprache als zentrales Koordinierungsinstrument. Sprache ist ein Interface zwischen den Wortsymbolen einerseits und der Realität andererseits. Dadurch kann die komplexe Realität in wenige Wörter komprimiert werden.

Erst durch diese Realitätskompression ist es möglich den Handlungsraum für Roboter zu verkleinern. Der Roboter plant nicht länger in einem 3d Raum mit Trajektorien sondern der Roboter plant mit Hilfe von Verben und Substantiven in einem abstrakten Sprachraum. Seit der Erfindung von Word embeddings wie Word2vec mag dieser Ansatz selbstverständlich klingen aber bis zum Jahr 2000 war der Fokus auf Sprache eine Revolution.

Noch immer steht natürliche Sprache ein wenig außerhalb der klassischen Computerwissenschaft. Es hat nichts zu tun mit Elektrotechnik, Mathematik oder Algorithmen, sondern die Sprachwissenschaft hat ihren Ursprung in den Geisteswissenschaften. Nicht nur in der Dewey Dezimalklassifikation wie sie in Bibliotheken verwendet wird, sondern auch in der Gliederung von Universitäten sind Geistes- und Naturwissenschaften unversöhnliche Gegensätze die getrennt betrachtet werden.

July 22, 2026

Natural Language as AI technology

Languages likes English or French are discussed by Linguists not by computer scientists. A language is located in the humanities but not within the mathematics department. So its logical that most Linguists have no idea about computer science and vice versa. This might explain why AI research wasn't succesful over decades, because Natural langauge is the missing puzzle piece to make machines intelligent.

From a birds eye perspective, a language like English consists of verbs, nouns and adjectives. The grammar consists of rules how to connect the words to sentences. And language also is directed towards the reality. A word like "a flying bird" or "green flower" is referencing to objects which are available in front of the speaker.

The ability to name every object from the reality with a word and describe activities also with words makes natural language a powerful technology which can be used for human to human communication and machine to human communication both. In case of Artificial Intelligence a computer needs to parse natural language which is the bottleneck in modern AI research. Suppose a computer program asks the user to enter a word. The user enters "red box", then the computer stores the input in a variable but it has no consequences. So the computer isn't able to understand the meaning.

This parsing problem can be solved by inventing a language game. A language game is similar to a 2d arcade game a rule book which explain who to react to a certain input pattern. Typical language oriented games are translation games, the board game Scrabble or the "guess what" game. After implementing these games on a computer, the input of a human will have consequences. The consequences are given by the rules of a certain game, for example in a translation game the user needs to enter the correct translation for a word from another language. The computer verifies if the answer is correct.

Strictly spoken, not the computer decides about the meaning but the rules of the language games are providing the meaning. Modern AI research after the year 2020 is mostly focused on natural language and language games like "Visual question answering", "instruction following" and "question answering chatbots". All these games are formalizing the human to machine interaction in a sense that the computer can determine a score. This score allows to train artificial neural networks. The result of the training process is recognized as Artificial Intelligence in the modern sense. It allows to control robots with language instructions, and generate text with Large language models.

A common assumption in the past was that its very complicated to parse natural language with a computer. This assumption is only correct if the entire corpus of English should be understand by a computer which is around 1 million words and stored in endless amount of books. Parsing this written language with a computer is indeed a hard problem for computer science. What is possible instead is to reduce the task to a subset of English which consists of a dozens words from a restricted domain which are used to play a language game. In the minimal case, there are 6 picture cards and 6 word cards and the task is to match the correct pairs. Such a language game can be implemented in a short computer program and can be played by an automated AI algorithm. The AI algorithm has access to a database with the correct answers. This allows the computer to connect the picture of a banana with the word "banana".

June 19, 2026

Sprachverstehen durch Computer

Zuverlässige Spracherkennung funktioniert nur in Science fiction Serien aber nicht in der Realität. Über Jahrzehnte war es ein ungelöstes Problem der Informatik ein natürlich-sprachliches Interface zu programmieren. Mit ein Grund dürfte darin liegen, dass aus Linguistischer Perspektive unklar war, was genau natürliche Sprache eigentlich ist.

Man kann Sätze als String-array in Computern speichern und sogar Subjekt / Verb und Objekt erkennen, nur folgt daraus nichts für einen Computer. Ein Computer versteht nur eine Sprache und das ist Assemblersprache oder notfalls eine Programmiersprache wie C/C++. Natürliche Sprache funktioniert nach komplett anderen Regeln. Um den Gap zu schließen gitl es das Problem Spracherkennung zunächst einmal mathematisch zu beschreiben in Form eines Datasets. In der ersten Spalte werden natürlich sprachliche Kommandos abgelegt wie "fahre zum Regal B" während in der zweiten Spalte eine Sequenz von Bildern hinterlegt ist die Zeigen was der Roboter tun soll.

Dieser Dataset definiert was das Problem ist und zwar soll der Computer so agieren wie in dem Dataset dargestellt. Erst in einem zweiten Schritt überlegt man sich dafür passende Alogirthmen oder entwirft neuronale Netze welche die Fehlerzahl möglichst minimieren. Sprachverstehen ist nach dieser Definition also die Fähigkeiten einen vorhandenen Dataset zu imitieren. Zuerst entwirft man einen Sprachtest und dann ermittelt man die punktzahl eines Computerprograms um diesen Test zu bestehen. Das ist das Grundprinzip beim Deep Learning wie es seit den 2010er Jahren erfolgreich in der Informatik erforscht wird.

March 23, 2026

Language enabled artificial intelligence

AI resarch in the past was dominated by an algorithm centric bias. The goal was mostly to invent an advanced computer program which simulates intelligence. The idea was inspired by a chess engine which is searching in the game state for the next action. Robotics projects were engineered with the same objective. Notable algorithms are RRT for kinodynimaic planning or genetic algorithm for artificial life.

In the 1990s and until 2000s this paradigm was accepted as state of the art attempt to realize artificial intelligence. Unfortunately. none of the described techniques was successful. The robots are not working and the Artificial life simulation didn't evolve into a life form.

There was something missing until the 2000s for enabling artificial intelligence, and the missing element is natural language. On the first look this explanation doesn't make sense because there are lots of examples for text adventures and language understanding AI projects in the past, so the principle isn't new and can't be the explanation how to realize AI. Typical well known examples from the past are SHDRLU, the Maniac mansion game which was based on a simple 2 word 'English parser and a speech enabled robot from 1989 (SAM by Michael Brown). 

The breakthrough technology after 2010 was to focus again on language guided robotics and implement these projects with more effort. Instead of programming a video game like Maniac mansion the goal was to program a text interface for robot control. Instead of realizing the parser in a simple C code, the parser is realized with a neural network. So we can say, that AI after the year 2010 has put natural language into the center of attention and created new algorithms and software around the problem. This attempt was very successful. It is possible to control robots with language and very important its even possible to scale up the approach so that the robot will understand more words and solve more demanding tasks.

From a birds eye perspective the situation until 2010 was to implement closed systems. AI was imagined as an autonomous system which is operating with algorithms and has no need to talk to its environment. In contrast, AI after the year 2010 works as open system. The robot receives commands from the human operator and sends signals back to the operator. In other words, modern robotics is always remote controlled with a text based interface. Its not possible to implement Artificial intelligence somehow else, but text based interaction is the core element of any AI system.

Its only a detail question how to program such an open system in detail. One attempt might be to utilize neural networks and learn the language from a dataset. Another attempt is to program the parser in a classical computer program without neural networks, while the third technique is to invent a domain specific language used for human to robot interaction. All these approaches have in common that natural language is the core building block. Natural language is used as an abstraction mechanism to compress the complex reality into a list of words. A typical entry level robot knows around 200 words to describe the environment including possible actions. These words are used by the robot to interact with a human operator. So the AI problem is mostly a communication problem, similar to transmit messages over a wire.

The paradigm shift from former computation into modern communication is the breakthrough technology for enabling artificial intelligence. Natural language used for human communication is also a powerful tool for human to robot communication. The English vocabulary is seen as a hammer for solving problems. Any problem in robotics gets reformulated into a language problem. There is no limit visible, but all the problems like biped walking, navigation in a maze and pick&place can be reformulated as language games.

Such kind of utilization of natural language wasn't available before the year 2010. The only thing known were isolated projects which explored if language might be useful for robotics. There was no understanding available that natural language is the core element in AI and needs to implemented in any possible robot or AI problem.

March 08, 2026

Die Technisierung der Sprache

Die Geschichte des Computers war anfangs eine Automatisierung des rechnens. Der abakus war ein frühes Beispiel für den Einsatz von Hilfsmitteln, während ab dem 17. Jahrhundert mechanische Rechenmaschinen erfunden wurden. Später mündete die Entwicklung in Tabelliermaschinen sowie der Entwicklung des elektronischen Computers, der Harvard Mark I. Computing bedeutet übersetzt soviel wie Rechnen mit Maschinen.

Es verwundert nicht dass die Informatik aus der experimentellen Mathematik entstand. Frühe Rechengeräte waren Hilfsmittel der Mathematiker um logarithmen zu berechnen oder numerische Simulationen durchzuführen. 

Die Geschichte der rechnenen Maschine kann als voller Erfolg bezeichnet werden, über einen Zeitraum von 300 Jahren wurden die Apparate kontiniiurlich verbessert und heutige Computer verbrauchen sehr wenig strom können aber sehr schnell rechnen.

Leider ist diese Fähigkeit für den Bereich der künstlichen Intelligenz uninteressant. Das rechnen mit Zahlen allein ist nicht ausreichend um Probleme zu lösen. Was es hierfür braucht ist natürliche Sprache die mittels computer verarbeitet wird. Es gab innerhalb der technikgeschichte durchaus versuche Computer zur Sprachverarbeitung einzusetzen, allerdings ging die entwicklung nur langsam voran.

Es dauerte beispielsweise bis in Jahr 2010 bevor der erste Sprachgesteuerte Gabelstapler vorgestellt wurde, der M.I.T. Forklift. Und das obwohl die KI Forschung sowie die Informatik-Community als sehr innovativ bezeichnet wird. Scheinbar gab es die Jahrzehnte davor andere Prioritäten als sich ausgrechnet mit Mensch zu Maschine Kommunikation zu beschäftigen.

Doch zurück zu klassischen Computern welche nur Zahlen aufaddieren können. Zahlen werden durch die Mathematik verwendet um Dinge zu beschreiben, z.B. die Koordinaten in einem 2d Raum, oder die Schwingung eines mechanischen Pendels. Diese Reduktion auf den numerischen Raum hat vorteile aber auch Nachteile. Der vorteil ist, dass sich darüber die Wirklichkeit vereinfachen lässt, ein weiter Vorteil ist, dass mathematische Gleichungen sehr gut von Computern gelöst werden können. Als großer Nachteil ist zu nennen, dass Zahlen allein nicht ausreichen um KI Probleme zu lösen. Insbesondere zur Beschreibung von Mensch zu Roboter Kommunikation braucht es neben Zahlen noch weitere Zeichen, allen voran Codes, Wörter und Sprachen.

Die details dieser nicht-zahlenmä0igen Datenverarbeitung wurden erst relativ spät in der Informatik als Problem erkannt und folgerichtig näher untersucht. Der Einsatz von Computer zur Verarbeitung von Sprache war kein einmaliges Ereignis sondern es gibt sehr viele Projekte dazu aus unterschiedlichen Dekaden. Das Leuchtturm projekt war zweifelsfrei SHDRLU (1968), aber auch das IBM Watson Q&A System (2011) ist erwähnenswert.

Eine mögliche Erklärung für diese langsame Entwicklung von redegewandten Computern könnte etwas mit gesellschaftlichen Ängsten zu tun haben. Klassische zahlenverarbeitende Rechner sind breit akzeptiert. Geräte wie Taschenrechner oder Softwareprogramme die Zahlenkollonnen aufaddieren sind weit verbreitet und niemand fühlt sich davon herausgefordert. Anders sieht es bei Software aus, die Sprache verarbeiten kann. Projekte wie der Eliza chatbot von Weizenbaum, der Word2vec Algorithmus oder das Zork I textadventure gelten als Blasphemie. Also etwas das mit dem Teufel im Bunde ist und das man besser meidet ...

March 07, 2026

Natural Language is a technology

Classical examples of technologies are machines like a steam engine or a mechanical calculator. These artifacts are presented in a museum about history of technology and its evolution is explained. There is a much bigger technology which is usually ignored by the classical research called natural language.

Natural language is a communication code used by humans to send messages back and forth. Typical languages are German, English or Mandarin. The reason why languages are usually not recognized as technology, is because its not a machine but used entirely by humans. During oral interaction there is no need for external apparatus because language can be generated only with the own voice and is received by the biological ears.

The reason why it makes sense to treat language as technology too is because it had a great influence for the evolution of technology including machines. Milestone technology in the history of mankind like the Gutenberg printing press, the telegraph or the internet have the purpose to improve communication by language. For example, the HTML file format allows to store large amount of written text in a digital format which can be books, online forums and newspapers. The Gutenberg printing press was invented for a similar need. It allows to duplicate the written word.

Modern examples for language processing machines are large language models which are used in chatbots and voice controlled robots like the figure.01 humanoid robot. In both cases, the machine is processing natural language with algorithms which results into a very powerful technology.

The assumption is, that future machines, not invented yet, have also the main purpose to process natural language. Its likely that artificial intelligence is realized by advanced language understanding algorithms. This allows to mechanize words, phrases and communication in general.

timeline:

1440,printing press by Johannes Gutenberg
1844,morse code by Samuel Morse
1915,Therblig notation by Frank Gilbreth
1954,Georgetown-IBM Experiment with russian translation
1968,SHRDLU natural language understanding by Terry Winograd
1989,Speech activated manipulator SAM by Michael Brown
2003,M.I.T. Ripley robot by Deb Roy
2011,IBM Watson Question answering by David Ferrucci
2023,Wayve Lingo-1 self driving car 

January 23, 2026

Creating an internal teacher with natural language

 Natural language is a powerful tool for humans to describe the reality. THe existing vocabulary can be reused for robotcs applications. The only bottleneck is, that computers can't parse natural language directly but need a parser and an ontology for doing so. The following blog post explains the idea for a kitchen robot.

The starting point is an ontology which is realized as a python dictionary. The ontology stores items in a kitchen and possible actions for these items.

--------------
items:
  apple, food
  banana, food
  plate, dishes
  pot, dishes
action
  prepate_meal, search(food)+eat(food)
  cleanup_kitchen, search(dishes)+search(trash)
--------------

The system is realized as a command line prompt which asks the human operator to enter commands. A possible session is shown next:
$ apple
> apple is food, location is (10,1)
$ milk
> not_found
$ banana
> banana is a food, location is (20,4)
$ prepare_meal
> search(food) -> found: apple,banana
> eat(food) -> eat(apple), eat(banana)

The parser takes the human input and searches in the ontology for a definition, for location and for actions. Its some sort of text adventure game which works also with a parser and a database. If the user enters a command or object name which is not available in the database, the parser will answer the request with an error message. In other words, the intelligence of the AI doesn't depend on the parser itself but its the result of well populated database. Somebody has enter all the kitchen items into the database to make the system highly responsive.
  
  

July 20, 2025

Simple chatbot in python

 The most basic implementation of a chatbot works with predefined questio answer pairs stored in a Python dictionary. The human user has to enter exactly the predefined question to get an answer so the chatbot is a database lookup tool. Even if the program is less advanced than current large language models and its less mature than the Eliza software, its a good starting point to become familiar with chatbot development from scratch. The sourcecode consists of less than 50 lines of code in Python including the dataset.

import re

def run_chatbot():
    knowledge_base = {
        "hello": "Hi there! How can I help you today?",
        "how are you": "I'm a computer program, so I don't have feelings, but thanks for asking!",
        "what is your name": "I am a simple chatbot.",
        "who created you": "I was created by a programmer.",
        "what can you do": "I can answer questions based on my internal knowledge base.",
        "tell me a joke": "Why don't scientists trust atoms? Because they make up everything!",
        "what is the capital of france": "The capital of France is Paris.",
        "what is the largest ocean": "The Pacific Ocean is the largest ocean.",
        "what is the highest mountain": "Mount Everest is the highest mountain in the world.",
        "what is the square root of 9": "The square root of 9 is 3.",
        "what is the weather like today": "I'm sorry, I cannot provide real-time weather information.",
        "how old are you": "I don't have an age in the human sense.",
        "what is python": "Python is a high-level, interpreted programming language.",
        "what is AI": "AI stands for Artificial Intelligence, which is the simulation of human intelligence processes by machines.",
        "where are you from": "I exist in the digital realm!",
        "can you learn": "I don't learn in the same way humans do. My responses are pre-programmed.",
        "what is gravity": "Gravity is a fundamental force of nature that attracts any two objects with mass.",
        "what is photosynthesis": "Photosynthesis is the process used by plants, algae, and cyanobacteria to convert light energy into chemical energy.",
        "what is the speed of light": "The speed of light in a vacuum is approximately 299,792,458 meters per second.",
        "thank you": "You're welcome! Is there anything else I can assist you with?"
    }
    print("Welcome to the simple Q&A Chatbot!")
    print("Type 'quit' or 'exit' to end the conversation.")
    print("-" * 40)

    while True:
        user_input = input("You: ").strip().lower()

        if user_input in ["quit", "exit"]:
            print("Chatbot: Goodbye! Have a great day.")
            break

        found_answer = False
        for question, answer in knowledge_base.items():
            if question in user_input:
                print(f"Chatbot: {answer}")
                found_answer = True
                break
        if not found_answer:
            print("Chatbot: I'm sorry, I don't understand that question. Can you please rephrase it?")

if __name__ == "__main__":
    run_chatbot()

January 08, 2025

Von pseudo Robotern zu echten Robotern

 

Wie man einen Roboterarm fernsteuert ist seit mindestens den 1980er Jahren bekannt. Ein menschlicher Bediener bewegt einen Joystick, dessen Signale werden über ein Kabel an einen Robterarm übertragen und dieser bewegt sich dann. Der Grund warum diese technische Aparatur nicht als echter Roboter gilt ist weil der menschliche Bediener die ganze Zeit über anwesend sein muss und auch nur den einen Roboterarm zeitgleich steuern kann, aber nicht 2 oder noch mehr.
Um echte Robotik zu realisieren, die ohne menschlichen Bediener auskommt, muss man etwas tiefer in den Methodenkoffer der Informatik hineingreifen und dort nach Sprachinterface Ausschau halten. Insbesondere in ihrer Schriftlichen Form kann damit Automatisierung erzielen. Man notiert die Befehlssequenz in einer Textdatei ähnlich wie ein Computer program und führt diese dann aus. Während der Ausführung muss kein menschlicher Bediener anwesend sein. Je abstrakter die Sprache gewählt wurde, desto einfacher ist die Steuerung der Robotern. Man notiert in dem Program beispielsweise: “greife den Apfel, das wiederhole dann 100x” und schon hat man einen Subtask automatisiert.
Der Grund warum dieses Prinzip bis ungefähr 2010 nicht in der Robotik eingesetzt wurde hat einen simplen Grund. Es gab früher keine Sprachinterfaces die leistungsfähig genug waren. Es gab zwar einige Experimente wo über Sprachkommandos Roboter gesteuert wurden, aber es war unklar wozu das nützlich sein könnte.
Die Sprachsteuerung von Robotern hat einen winzigen Nachteil: und zwar werden Roboter dadurch sehr menschlich. Es sind keine simplen Automaten oder Rechenmaschinen mehr die elektrische Impulse weiterleiten und Zahlen aufaddieren sondern künftige Roboter werden die solben Worte verstehen wie Menschen auch, also Vokabeln wie “schnell, stop, links, rechts, greife, Apfel, Banane, Tisch” usw. Dadurch geht der Unterschied zwischen Mensch und seiner Technologie verloren. Anders als die Fähigkeit eine mathematische Rechnung auszuführen ist die Fähigkeit Substantive und Verben zu verstehen ein sicheres Zeichen für intelligenz. Computer wie sie früher verbreiteet waren waren dazu nicht imstande. Es waren schnellere Taschenrechner die über eingebaute Gleitkommaarithmetik verfügten aber nicht über Sprachprozessoren verfügten.
Es gibt einen großen Unterschied zwischen menschlicher und maschineller Sprachverarbeitung. Wenn Computer Sätze parsen geht das tausendmal schneller. Ein Computer muss nicht 2 Wochen einplanen um einen Roman zu lesen sondern Computer erledigen diese Aufgabe in unter 1 Minute. Diese messbare Beschleunigung ist ein sicheres Anzeichen für technischen Fortschritt weil es bedeutet, dass Menschen ihre wichtigste Fähigkeit automatisiert haben.

September 30, 2024

Chatbot Evolution in den 2010er Jahren

 Es ist naturgemäß schwierig aktuelle Entwicklungen unter historischen Aspekten zu beleuchten, weil die Menge an Literatur zunimmt je näher man sich der Gegenwart nähert und weil die Strömungen schwerer zu überblicken sind. Anstatt die tatsächliche Gegenwart in Bezug auf Chatbots zu beschreiben welche von 2020 bis heute geht, besteht eine mögliche Alternative darin, etwas weiter zurück zu gehen und nur die 2010'er Jahre zu beschreiben.

Einerseits gab es in den 2010er neu entwickelte Algorithmen im Bereich Machine Learning und Natura language processing um Chatbots an sich zu verbessern.  Es gab darüberhinaus aber noch eine weitere weit weniger offensichtliche Technik und zwar die Einführung von Chatbots benchmarks. Dabei geht es nicht darum, einen chatbot in der Leistung zu erhöhen ihn also menschlicher zu gestalten sondern bei einer chatbot challenge geht es darum, vorhandene Chatbots untereinander in ihrer Leistung nach Punkten zu bewerten. Die Annahme lautet dass jemand anderes bereits mehrere Chatbots programmiert hat und es darum geht diese zu ranken.

Die Entwicklung dieser Chatbot vergleichsbenchmarks sind der eigentliche Grund der Verbesserung der Chatbot technologie. Bevor man neuartige Algorithmen inkl. neuronaler Netze entwickelt muss man zuerst einmal wissen wie vorhandene Konzepte leistungsmäßig abschneiden. Praktisch werden die Benchmarks als Dataset realisiert, was übersetzt soviel wie Datenbasis oder Tabelle bedeutet.  Datasets haben ihren Ursprung im Machine learning wo man zwischen Training dataset und test dataset unterscheidet.

Der Übergang von manuell programmierten Chatbots hin zu Dataset benchmark wird durch den ALICE chatbot (1995) aufgezeigt, der das AIML format verwendete. AIML ist einerseits das Datenformat für den Chatbot aber dient gleichzeitig als Korpus für eine Wissensdatenbank. Wenn man jetzt nur den Datensatz verbessert aber nicht den Chatbot, erhält man einen Chatbot dataset. Wo also die Entwicklung eines Punktesystems im Vordergrund stteht. Andere chatbots haben nun die Aufgabe, in Bezug auf einen konkreten AIML Korpus eine möglichst hohe Punktzahl zu erzielen.

Spätere Chatbot benchmarks basierten nicht länger auf dem AIML format sondern verwendeten csv dateien oder sogar Textdateien. Der Benchmark wurde in Form einer dokumentensammlung bereitgestellt was den Übergang zu Question & Answering systemen darstellt. Die Aufgabe für diese Generation von chatbots bestand darin, Fragen zu einem Dokumentencluster zu beantworten. Auch hierbei gab es Algorithmen, die dabei sehr gut abschneiden und andere denen es weniger gut gelang.

September 29, 2024

Chatbots bis zum Jahr 2010

 Die Geschichte von Chatbots kann man grob in 2 Phasen einteilen: 1960-2010 und 2010-heute. Die ältere Zeitperiode ist gut dokumentiert und leicht nachvollziehbar. Projekte wie Eliza und Cleverbot wurden programmiert um menschliche Dialoge zu imitieren. Sie basieren auf einem parser, der die Eingabe des menschlichen Benutzers auswertet und einem Satzgenerator der daraufhin eine Antwort erzeugt. In der Interaktion führt das dazu, dass Chatbots einerseits Sätze erzeugen aber gleichzeitig die Schwächen der KI deutlich werden.

Ab Anfang der 2000er Jahre gab es in der Chatbot Technologie eine Neuerung und zwar das AIML Dateiformat. Darin können Frage / Antwort Paare außerhalb des eigentlichen Quellcodes abgelegt werden. Man muss also einerseits die Chatbot software programmieren und dann noch eine AIML Wissensdatenbank erstellen die eine normale Datenbank ist. Aber, auch mit AIML ist die Leistungsfähigkeit der erzeugten Chatbots nicht besonders hoch. Es ist nur leichter diese zu programmieren.

Ab dem Jahr 2010 gab es in der Chatbot Entwicklung eine massive Verbesserung. Diese hatte nur indirekt etwas mit Maschine Learning zu tun sondern die Veränderung bestand darin, zunächst einmal das Problem zu definieren um das es geht. Bis 2010 war die implizite Annahme es geht darum einen Chatbot zu programmieren, also eine Computersoftware die auf Sprache reagiert. Nur, das ist nicht das eigentliche Problem. Worum es wirklich geht ist das Question answering Problem zu lösen. Als Q&A challenge bzw. als Q&A dataset wird eine Tabelle mit Quizfragen bezeichnet die ein Mensch oder ein Computer richtig beantworten muss. Es geht also eben nicht um die innere Funktionsweise einer künstliche Intelligenz sondern es geht um eine Wissensspiel ähnlich wie Trivial Pursuit, Jeopardy usw.

Bevor man einen Chatbot programmieren kann muss man zunächst einmal eine Aufgabe programmieren die dieser Bot lösen soll. Also eine Art von Computerspiel ähnlich wie Tetris, Pong usw. Allerdings geht es bei diesem Spiel anders als bei Arcade spielen nicht um schnelle Reaktionsfähigkeiten sondern um das Verstehen von Sprache. Daher auch die Bezeichnung Question&Answering challenge. Dieses Spiel funktioniert so: dem Spieler wird eine Frage präsentiert, z.B. "Was ist 5+2?" und der Spieler muss antworten "7". Eine weitere Frage könnte lauten "Welcher Kontient ist der trockenste und heißeste der Welt?" Antwort: Afrika.

Für ein Quizspiel spielt es keine Rolle ob jemand darin gut abschneidet oder schlecht. Und ob jemand das Spiel als Mensch löst, als Computerprogram oder als neuronales Netz. Sondern ein Quizspiel ist nur ein formalisiertes Problem. Es besteht aus einer Anzahl von quizfragen aus unterschiedlichen Bereichen und es gibt Antworten auf jede Frage die geheim sind für die Kandidaten. Jetzt kann der Quizmaster ermitteln wieviele Antworten ein Kandidat richtig beantwortet.

Das besondere an Q&A challenges ist dass man damit die Performance von Computersystemen bestimmen kann. Man kann unterschiedliche AI Algorithmen entwerfen die die Fragen beantworten und dann lässst sich sagen, welcher Ansatz besser war. Q&A challenges sind die Grundlage um fortschrittliche Chatbots zu entwickeln. Ein chatbot ist schlichtweg eine Software, die in einer Q&A challenge besonders gut abschneidet.

Mit diesem Hintergrundwissen lässt sich besser herausarbeiten was der Unterschied ist zwischen den Chatbots bis 2010 und jenen die danach entwickelt wurden. Chatbots vor 2010 wie Eliza wurden nicht mit Punkten bewertet. Die Software hat zwar einen Dialog geführt aber es gab keine Punktzahl über die Qualität der Ausgaben. Bei neueren Bots wie IBM Watson gibt es eine solche Bewertung. Eben weil neuere Chatbots als Antwortgeneratoren für Q&A Challenges entwickelt wurden.

Um einen neuen Chatbot zu programmieren benötigt man keine besonderen Algorithmen aus dem Bereich sondern benötigt wird ein conversational dataset. Also eine Tablele mit Frage/Antwort Paaren aus dem Bereich smalltalk.  Dieser Dataset wird als Benchmark und als Problemdefinition verwendet. Der Dataset ist eine Art von Small talk spiel und im zweiten Schritt kann man ein Computerprogramm erstellen, dass dieses Spiel spielt. Im wesentlichen geht es also darum zu unterscheiden zwischen einer Aufgabenstellung (dem Q&A Dataset) einerseits und einem Teilnehmer (dem chatbot) der innerhalb des Spiels aktionen ausführt.

March 25, 2023

The difference between a pocket calculator and GPT-3 enabled Artificial Intelligence

In contrast to AI based software a pocket calculator was understood by the public quite well. Such technology is available at least since the 1970s and lots of books are explaining the inner working. The core element is a CPU which gets programmed in Assembly. And then the machine allows the user to enter a task like “2+4”. From a technical perspective a pocket calculator contains of the CPU which is in easiest case an 8bit model, there is some sort of onboard RAM and very important a program which takes the input of the user and sends instructions to the CPU.
Creating yet another pocket calculator in hardware is easy. Also it is possible to write a python software which emulates a pocket calculator. There are endless amount of software and hardware components available for this purpose and they can be explained easily to newbies. It is more complicated than normal mathematical but the technology is not very advanced.
In contrast, a modern gpt-3 driven Artificial Intelligence works very different to a pocket calculator. Even the system contains of hardware and software components it can't be grasped with the traditional terms used in computer science. Also it is more complicated to newbies what a neural network is about. The inner working of AI can be summarized the following way.
The AI software is able to convert back and forth from natural language into numerical arrays. The input of the user gets converted into numbers, then the system is doing something with the numbers, and the output is converted back into natural language. This number-word engine is the core element in Artificial Intelligence. it allows to solve any problem. Over decades it was unknown how to do so, and it was even unclear if such a transition is needed. Creating word embeddings is sometimes called the symbol grounding problem. For example the word cat is not only a sequence of single characters (C + A + T) but cat is represented in a conceptual space as a number next to other words like dog, mice and so on.
The surprising situation is, that after solving the word embedding problem, it is quite easy to construct a human level Artificial Intelligence. If a problem was reformulated as a numerical mathematical problem, existing computer technology can be applied to it which means the information are stored in the main memory and there are routines which are doing something. The only bottleneck is the transformation back and forth from words to numbers. The core element of any advanced Artificial Intelligence which allows to rebuild the gpt-3 software from scratch is a word embeddings algorithm. Such a software component takes an English sentence and converts it into a mathematical vector. The details of this sophisticated technology are not understood very well and it is a very new approach.

January 27, 2023

A rough description of the perplexity.ai chatbot

 

In contrast to the famous chatgpt software, the perplexity.ai chat is less known. It's main advantage is, that everybody can use it. Similar to chatgpt, it is difficult to get concrete information how the algorithm is working in the background. Somewhere it was mentioned, that the same gpt3 language model was used which says anything and nothing. The more interesting question is what are the limits of the chatbot? Is AI really new technology or is the hype around chatgpt exaggerated?
In contrast to a search engine like Google, the perplexity.ai software is a Q&A software and can answer queries formulated in natural English. To make the difference clear, we have to describe first what the limits of a normal search engine are. Suppose the user enters two and more keywords into the input field:
word1 AND word2 AND word3
In many cases the result list will become empty because a website contains only of two of the words, but not three of them. The google search engine can't guess what the user is asking for but the result page is generated with normal non AI technology.
In contrast, a Q&A software is converting the internet and the user request into a numerical vector and then searches for the best matching. So it is basically a neural network which understands English. This allows to parse and generate the information more robust.
Some years ago there was the “Wolfram Alpha” search engine which has promised to revolutionize online search but a detailed investigation has shown the limits of the project. So let us try what the perpelexity.ai is able to do.
The basic question for any Q&A software is a simple YES/NO question like, is the ocean blue, is the sun yellow and is pizza a food. The chatbot answers all of them correctly and gives smart additional information.
The next iteration is a more complicated request like “What is 10/3?” “Which books were written by Hemingway?” “When was the computer invented?” All these questions were answered by the chatbot with ease. The assumption is that the better known chatgpt is also capable of this sort of tasks.
So we can say, that the perplexity.ai chatbot is the first working q&a webseite which produces high quality output.
Let us take a look at the more serious part which is about possible failure of the software. What the perpelxity.ai software can't do is to write longer stories. All the output is limited to 3-5 sentences. Also the software can't tell a joke but it will provide only the URL to a website which hosts some jokes. In addition the chatbot can't display images and in most cases the answers explaining only the surface of a topic but not the detail information.
What perplexity.ai and chatgpt have demonstrated is, that from a technically side it is possible to program a Q&A website which is more powerful than a traditional search engine.
But let us go back to more basic tasks. The chatbot can list the major books written by hemmingway. In contrast to the google search engine it is not showing a website somewhere in the internet and it is not redirecting the user to Wikipedia which has the same list. But the chatbot prints out the information in his own words. For the untrained user this sounds like magic
To understand the difference to a chatbot from the past let us ask the same question to the Cleverbot chatbot:
Human: Which books were written by Hemingway?
Cleverbot: I don't remember.
Human: Again, Which books were written by Hemingway?
Cleverbot: What do you mean I was written?
Human: Hemingway list of books
Cleverbot: What is your favourite book?
Such kind of interaction was normal for pre chatgpt software. The software is using a predefined set of answers and can't answer concrete questions for facts.
In contrast, the perplexity.ai chat bot can parse the question and gives the correct answer. And it can so for a slightly different question as well.
This might explain why the chatgpt software has gained strong attention in the internet. Because the beta testers are surprised how powerful the software is and they want to know what the limits are. One possible answer is, that the engineers themself doesn't know the answer. A likely comparison is to describe chatgpt as some sort of chess playing engine which is able to beat the strongest human player in the world. It was difficult to accept for humans that computers can play better chess and some journalists have argued that the AI isn't understanding on a deep level what chess is about. Perhaps it is possible to give the same argument for modern Q&A website. On a technical level the machine can only process numbers but has no understanding what the meaning is.
The ability to process natural language is usually seen as a sign of intelligence. In most exams the students are asked to answer easier and more difficult questions formulated in English. In a typical exam about computer science a possible list of question would look like:
Who has invented the TeX software?
What is a c compiler?
Explain the advantages of object oriented programming.
How does the A* algorithm works.
Most students can't answer all these questions correct but only a fraction of them. The teacher will read the answer and determines the score and then the human students gets a certain review. Suppose a computer program is able all these question correctly, does this mean that the AI is on the same level like a human student? What we can say is, that AI is able to imitate humans in a surprisingly high level.

June 13, 2018

Natural language interface for solving the frame-problem with a layered game-engine

Title: Natural language interface for solving the frame-problem with a layered game-engine

Author: Manuel Rodriguez

Date: 13. June 2018

Abstract: A planning algorithm like A* and RRT works only efficient for small problem spaces. Most robotics problems have a huge problem space. The answer to the mismatch is a domain model which is enriched with heuristics. In the early history of AI this problem is discussed under the term “frame problem” and means to describe the action model in a object-oriented programming language. For the example of an textadventure and the parking with a car, a game-engine is presented which is using natural language commands for storing domain specific knowledge.

Table of Contents

1 AI Planning
1.1 Knowledge based planning
1.2 The GIVE challenge
1.3 Automatic creating of a pddl file
1.4 Theorem proving for beginners
1.5 Storing domain knowledge in a game-engine
1.6 Plan stack in HTN-planning
1.7 Building a simulator for HTN-planning
1.8 Combining autonomous and remote-control
2 Example
2.1 HTN Planner
2.2 Car physics
References

1 AI Planning

Mindmap

1.1 Knowledge based planning

Artificial Intelligence in simple games like chess and more complex robotics domains can be realized with AI-planning:

“Planning is the most generic AI technique to generate intelligent behaviour for virtual actors.“ [8]

How a planning system looks like for a chess-like game is already known. It is a brute-force search technique in the game-tree. More complex games can be planned too, but the algorithm is more complicated. A combination of a hierarchical planning system together with machine-readable knowledge representation is the standard procedure. It means to store the game on different layers in a formal model so that it can be searched in real time. Well known forms of storing knowledge are STRIPS and PDDL. Both are languages for formulating a symbolic game engine. They are used for the high-level-layer of a planning domain.

But let us go a step backwards. The only known algorithm for solving a planning task is a search in the game tree. The algorithm is called A* or in newer literature RRT. RRT alone is not able to solve complex tasks, it must be implemented on different layers at the same time. The layers are depended from the game, they are representing the domain specific knowledge.

The bottleneck is, that for most games, no machine readable domain description is available. For example in a pick&place task for a robot, it is unknown how the story looks like and what the outcome of an action is. The programmer is not searching for a solving-strategy, instead he is searching for the game. That means, he must enrich the given game with detailed rules. Only these rules can be solved by the planner.

[12] calls the process “domain knowledge engineering” and describes a software (GIPO) which makes it easier to create a planning game from scratch. On page 4 an example is given. The original game is the “docker worker robot” and the programmer must define potential states of the robot like “load”, “unload”, “at-base”, “busy” and so forth to enrich the game with knowledge. The result is a pddl file which can be used in a planning task.

LfD & HTN

A human operator has an implicit domain model, he knows that before he can grasp an object, first the gripper has to be opened. This domain model is not available in sourcecode and has to be programmed first. To overcome the gap the "learning from demonstration" is the right choice for constructing a "hierarchical task network" from scratch.[3] [6] The human operator has a GUI in which he records a manual motion, and from this demonstration a machine readable ontology is created. This task model can be used by the HTN-planner for handling the task autonomously.

The aim is to transfer the knowledge from the human operator to the agent. The knowledge is similar to a walk through tutorial for games. It is a description of how to play a certain game. In a hierarchical task network the knowledge is formalized like a symbolic game engine. That is a software-modul which can predict future game states, e.g. robot-grasp-object -> object-isin-hand Usually the description is based on natural language. Instead of using simple variables in the gameengine like:

bool a,b,c
int d,e,f

the variables have sounding names like “bool object-isin-hand true/false” or “int distance-between-object-and-gripper”. The reason is, that the domain model is not primarily programmed for a machine but as help for the software-engineering-process. Creating a domain model is not a mathematical algorithm like A* but a software-engineering-task like UML and agile development. The result of this step is not a PDDL file or executable code, it is a human-readable paper which is called the specification of the game. The algorithm doesn't learn by itself the game, instead a version control system like git is used to bring the domain model to the next version.

Grounding means to combine tracking control with natural language.[9] The idea is not only to describe a task with “subtask1”, “subtask2” and so on, but with the correct words e.g. “heating the water”, “fill the cup”. It is not possible to describe a domain without using natural language. From a machine perspective perhaps, because all the words have no meaning for the computer, they are only labels. But for maintaining a task model by human engineers they need natural language. That means, a task model is foremost a dictionary and not a mathematical algorithm.

Language model

Under the assumption that every task model is based on natural language the question to investigate is how does the language model will look like for a certain domain? The GIVE challenge is trying to answer this by generating natural language instructions to guide a human for doing a task.[2] The idea is, that on the first hand, a dictionary is coded in computercode and can be executed like a game-engine, and the output of the engine is used to solve a task.

Perhaps an example will make the point clear. In [7, page 6] is on top of the page the map of a game visible. It is a normal maze game, in which the player moves around and can press button. Below the image the pddl description is given, which can be seen as a natural language game engine. It provides commands like “move”, “turn-left” and “manipulate-button”. The pddl description contains of two important aspects:

• natural language words, for example “turn-left” instead of a   simple “action2”
• state-action-pairs, that means by activating an action, the   system is in a new state

The overall system can be seen as a living dictionary. It contains one the first hand, words and action-names, and they can be executed by the user which brings the system to a new state. The pddl file contains knowledge about the domain.

1.2 The GIVE challenge

The term “Generating Instructions in Virtual Environments” (GIVE) is used for describing a programming challenge with the aim to generate natural language for a domain.[5] The output itself is usually produced by a PDDL solver, that means for a given current / goal pair the solver is trying to find a plan through the domain. The solver needs as a precondition the PDDL domain model.

A more colloquial description of the challenge is to compare it with a textadventure game which is enriched by a 3d map on top of the GUI. It is comparable to the early Point&click adventures in the 1990s in which the human-operator has some actions like “go north”, and must reach a certain point in the map while is doing subtasks. The interesting aspect in the GIVE-challenge is, that from a graphical point of view, everything is minimalist instead the idea is the task model and it's potential application to Artificial Intelligence. Another important aspect in the challenge is, that natural language is in the center of focus. That means, the idea is not only to play a game by an agent, but instead the idea is to generate natural language which guides a human who is already capable of understanding English.

From a programming point of view, the GIVE challenge is one of easier tasks. It needs less energy to solve the challenge compared with Starcraft AI or Robocup. The task is not so easy as programming a normal 2d computer game, but it can be mastered by beginners in AI.

1.3 Automatic creating of a pddl file

A domain model consists of natural language. From a technical side, the domain model is stored in a symbolic planning language like PDDL, OPL or ABPL. The first one (PDDL) is a classical language, while the second others have object-oriented features. The easiest way for creating a domain model is program the domain-description by hand. It is the same technique like a textadventure is created. A more sophisticated idea is to create the language model automatically from event logs.[13]

The idea is to record all the event in a section called “agent memory” and then construct out of the relationships the pddl / ABPL description. That means, at the beginning the agent has no domain model, he has to build one from scratch while he gets new experiences. In the literature the term “action model learning” is used. An action model is a symbolic game-engine which can be stored in a PDDL file.

Instead of explaining how to realize such a system, at first we must describe how to evaluate an action model. At first, we need a working symbolic game engine, for example an instance of an textadventure. This game-engine produces a stream of natural language. The user can input text, and the dialogue system gives feedback also in natural language. On the right screen there is an empty prototype. The prototype has the obligation to emulate the working game-engine. That means, the prototype observers the events, stores everything and after a while he is acting in the same way. The goal is to reverse engineer a symbolic game-engine.

1.4 Theorem proving for beginners

In the early history of Artificial Intelligence, theorem proving played an important role. In the context of the STRIPS planning system, such systems were capable to prove mathematical questions. But what is theorem proving exactly? Why it is so hard?

Theorem proving means basically to create a puzzle, for example Rubik's cube, and search for a sequence. The cube has a starting pattern, certain operations are possible and the goal is to bring the system into a goal situation. For example, make all sides clear, or bring only one side into a healthy condition.

Mathematical theorem proving works with the same idea in mind. There is a starting equation, a number of allowed operations and a goal situation. Like in the Rubik's cube example the idea is to search for plan, if such a sequence was found, the theorem was proven. In reality, theorem proving is equal to game-playing. A game is system which has allowed moves and it is up to the player to decide which moves he want's to execute. Automatically theorem proving works surprisingly simple. The so called SAT solver is using brute-force-search and that's all. If the problem is small like in the rubik's cube example, the STRIPS program is successful, in much higher state-space for example “a theorem prover for chess” is much more difficult to realize, because the number of possible plans is higher.

In the historic paper [4] of 1971, Nilsson introduces the so called “frame problem”. This means basically, that the STRIPS language is only a simple planning language and has no object-oriented feature for describing more complex problems. More recent planning languages like “A Better Planning Language” (ABPL) can overcome the frame problem.

Example

From school mathematics there are equations known plus rules which can used onto these equations:

a+4=7
a+4=7 |-4
a=7-4
a=3

The starting situation was an equation, and we have applied an operator to it. After the action, the equation is in a new condition. In the example, we selected the action manually, but it is also possible to formulate the problem in the STRIPS language. Such feature is integrated in most computer-algebra systems. The so called solver is playing around with the equation to fulfill a certain condition.

The “frame problem”

With STRIPS is a powerful planning language available for proving any theorem. The problem is now: how to formulate a robotics-problem in the STRIPS syntax? The question is discussed in the literature as the frame problem, because the assumption is, that frame based aka object-oriented programming is part of the solution.

A more precise formulation of the problem is given under the term “General game playing”. The challenge is here to invent a game from scratch. That means, as input the system gets a plan trace of checkers, or a textadvanture, and the system is able to construct all the game-rules and codes them into the Game description language. The sad news is, that until now “General game playing” didn't work very well. It is not possible to construct real games from scratch. But that is not a real bottleneck, because it is always possible to manually program in Strips, GDL or any other language. It is not necessary to use automatic programming for realizing a robot. 

In reality, the frame problem is equal to a software-crisis. That describes a situation in which a demand for software is there but no sourcecode is available because of many reasons. The software crisis in the area of operating systems was solved with Open Source software, and the software crisis for the special domain of game-playing and planning languages will be solved with Open Science. For example, if programmer A describes in a paper an UML chart, a ontology and executable code for implementing a pick&place robot, than programmer B is able to reproduce the result and use this as a basis for a more sophisticated system.

The frame problem, the grounding problem and the general game playing challenge can all be solved with a better science communication which works manually and is working with Open Access papers and Open Source software.

1.5 Storing domain knowledge in a game-engine

Domain knowledge has to be stored in machine readable form. Most literature questioned how exactly the data-storage should be, for example in the PDDL format, in ontologies or in semantic networks. More important is the question how the result will look like if the domain knowledge is available. It will look like the game-engine of a textadventure. That means, it is possible to send a command to the engine, and the engine will output the future state.

A symbolic game engine which is controlled by natural language is able to predict future states. For example, a command like “grasp apple” results into the output “apple is in hand”. This logical reasoning has to implemented in a part of the software called game engine. A game engine can be programmed in a textadventure markup language, in Javascript and even with ontologies. In the easiest form a textadventure is programmed in Python with object-oriented feature. But in general it is only a minor problem, more important is the information that domain knowledge and a working game-engine is equal.

Let us describe what in the so called ICAPS conference is usually done. In most cases a domain like Blocksworld is converted into a pddl description. But that is not what the inner goal is. The more precise description is, that a textadventure for the blocksworld domain was created, and apart from PDDL this can be done in any other programming language. At the end, a game-engine must be implemented which can be filled with user commands.

But why is a textadventure so important, isn't it possible to write a normal game-engine with a graphical interface? The terms grounding means to connect natural language with actions. Grounding means, that the engine can parse a command like “grasp apple”. Every grounding results into a textadventure, because a textadventure is about the understanding of natural language. The only open question is, how to program such adventure for a certain domain.

Existing textadventure which were programmed since the 1980s contains domain knowledge in a machine readable form. They have a clean interface, it is possible to send a request to it, and the game-engine calculates the follow state. A textadventure can be seen as the inner core of a working robot control system. It is the part in which the domain knowledge is stored.

In the context of Artificial Intelligence the general name for domain knowledge which is stored in a textadventure is “dialogue system”. Dialogue system like TRAINS and FrOz are usually programmed to support the human player. They have a planning feature out of the box, that means, the engine was programmed with the goal to find a path through it by a solver.[1]

1.6 Plan stack in HTN-planning

[10] describes on page 8 the plan-stack of a HTN-planning system. A so called plan stack is a list of high-level commands:

1. task1
2. task2
3. task3

and so forth. The term “stack” is correct but it is a bit misleading, because the datastructure of storing a plan is not very important. Any other data structure for example a SQL-table, a csv-file and so on would also work great. The more important aspect is, that before building the plan stack the programmer needs to know the name of the tasks. In the cited example, the tasks are having to do with the soccer domain. Their names represents elements of the soccer game, for example, “pass-the-ball” or “shoot”. The knowledge is not stored in the HTN-planner directly, but in the domain-model which is used by the HTN-planner.

I would guess that the stack and the HTN-solver are the least important part of the AI-system. That means, to store a list of task-names in a computer-memory is a trivial programming task and can be implemented easily. The bottleneck in Hierarchical task networks is somewhere else. Like i mentioned above it is the domain-model. In general we can say, that the domain model is equal to a high-level game-engine. It is some kind of textadventure. The user has commands which he can send to the game-engine and the engine is executing them. The game-engine has the obligation to predict future game-states. That means, after the command “pass-the-ball”, the game engine changes the position of the ball to the receiver of the ball.

Basically spoken, the domain-model of a HTN-planner is equal to a computergame. Because of it's high-level-nature it is often realized as a textadventure but can have additional graphics to visualize the internal states. Let us go into the details for the Robocup case. A game for playing Robocup can be programmed on many levels. At first it is possible to play the game physical with real robots, while in the 3D Simulation league only a physics engine is used. In the 2D Simulation a different kind of physics engine (perhaps a 2d one) is in the loop, and it is also possible to simplify the engine further to a more abstract domain model, which is working without a realistic engine but only with a idealized physics engine.

For example, it is possible to use the SFML graphic library to program a soccer game which is not very accurate. The game is not realistic but is similar to a Pacman clone. That means, the players have limited actions possibilities and the simulation would look like a Atari 2600 game. The trick with a htn-planner is, to combine different game-engines in layers. For example on top a high-level textadventure, in the middle a graphical representation and on the bottom a 3d realistics physics engine. Optimizing these different layers is the key factor for a successful HTN-planner.

Model acquisition

Surprisingly many papers are discussing the automatic creating of action models for HTN-planning. The idea is to use plan traces and a very complicated algorithm which generates the action model. To be honest, automatic generation isn't working. If the aim is to get a real system for example to play a game, only handcrafted models can be used. Realizing an action model for a HTN-planner has to be done as programming in general. That means, that lots of man years has to be invested, and some kind of version control system is needed. That means, the action model can not be generated autonomously, it has to be programmed by man.

There are many working examples available for example from the gaming-industry or from the Robocup challenge. In all cases the workflow was hand-crafted. That means, a team of 10 programmers has written the documentation, they have painted UML charts and they have implemented the action model in a simulator. The process can be called a “Software engineering task”. Plan traces may be useful but there is no algorithm but only a project which can transform them into a executable action model.

What can be done automatically is to use a given action model with an automatic solver. If it is clear that the task contains the subtasks “pass” and “shoot” and if it is clear, what the follow state is, that an automatic solver is able to generate the best plan. It is the same strategy which is used in computer chess to generate the next move. It is done entirely by the computer and human intervention is not needed.

What is possible, is to use an existing serious game (which has an API) and run on top a HTN-planner. If the game contains the action model, the tasks and the events it is possible to calculate the plan. The programming effort is only minimal in such cases. But, here the programming effort has to be invested by somebody else who has programmed the original game. That means, the action model is not generated from scratch it was programmed to in an external software engineering project.

1.7 Building a simulator for HTN-planning

HTN-Planning itself is easy to understand: a given “action model” is sampled by a solver and the found plan brings the system into the goal state. The more demanding task is to create the action model. Let us describe in detail how does it look like.

On a programming level every “Action model” is realized with an ontology. That means there are some C++ classes which are containing methods and attributes. They can call each other. The purpose of the ontology is to realize a simulator, that is a piece of software which is mimicry the reality. A well known simulator which is used in the Robocup domain is a physics-realistic simulator. It is realized with a dedicated physics engine and calculates what will happen if the player kicks the ball. But a physics-simulation is not enough, in the context of hierarchical task networks, many layers of simulators are combined together. There is a need also for a high-level-simulator, a tactical simulator and so on. All of these software moduls can be realized with ontologies, better known as UML classes. The question is only how to transfer a certain domain for example a dexterous grasping task into a simulator.

Again, let us imagine what the benefits are. Suppose we have a layered simulator for a dexterous grasping task. On the lowlevel it is a physics engine, on the mid-level a simplified 2d game and on the high-level layer a textadventure which can be controlled in natural language. If such a detailed model is available, it is very easy to solve the game. All what the HTN-planner has to do is generate random actions on different levels and search for a certain situation which is called the goal. The only open question is, that for most domains, such a layered simulator isn't available. And without such system, the HTN-planner has nothing to do.

The open question is: how to program for a certain domain, a layered simulator in an object oriented programming language. Programming only a realistic physics engine is not enough, what a HTN-solver needs is a mixture of different simulators because this will speed up the search process. Let us make an example for the Robocup domain.

In the high-level-simulator we can execute a command like “move-to-ball”. After executing the command, the engine places the player direct to the ball. The cpu-consumption for doing so is exactly zero. That means, the game engine simply moves the player position and that is. There is no collision detection, no path planning and no check if the battery is empty. On a second layer, this command is resolved into detail commands which are more realistic and only on the lowlevel layer a realistic 3d physics engine which includes collision detection is needed. Such a layered ontology based simulator can be called the main part of a HTN-planner. It is a useful tool for generating a plan for an agent.

layered abstraction

Because the layer architecture is fundamental for HTN-planning, I want to a give an example. Let us suppose we are again in the Robocup domain. A normal simulator would implement a physics engine in 3D. In HTN-planning this is not enough. The idea is to search for a plan on different abstraction levels and especially on layers which are computation inexpensive. A 3d accurate physics-simulation is very cpu-demanding, it is the worst choice for testing out random plans in it. The better approach is to construct first an artificial game on top of the simulator. We can call this a simplified soccer simulator. It is only 2d based and has no dedicated physics engine. Instead it works like a board game. That means, the pieces on the table can be moved freely without any constraint. If agent #1 should go to a certain place we can manipulate his position directly. That means, the action “move” needs no cpu-consumption instead it is executed directly. On this simplified board game, it is possible to figure out basic strategy. For example, if we want that 3 players are on the same place, we must move all them to that place. The abstract game engine can answer certain question. All of them are abstract nature. It is a game on top of the original game.

Now we can combine both layers together. On top the 2d abstract game with a simplified mechanic and on bottom the elaborated 3d simulator which simulates an accurate physics engine. Before we will ever try out a move in the 3d simulator we are asking first the high-level-layer what the result will be. The overall domain must be implemented in at least 2 layers of different game engines. That is the basic concept of an action model behind a HTN-planner.

The high-level layer didn't replace the former 3d physics engine. An accurate physics engine is great for planning lowlevel actions, for example if we need to know how strong we must kick the ball until he reaches a certain velocity there is no alternative to a physics engine. It gives the exact result back. But, only a small amount of question of an autonomous agent has to deal with this detail question. In many other cases the agent needs advice if he should kick the ball to position A or B, no matter how exactly this can be done. This high-level-decision can't be answered by an accurate physics engine.

In reality, a domain is modelled by different game engines which are arranged in layers. At least 2 layer are needed, but more layers are better. This raises the problem of how exactly the different layer-engines has to be implemented. Manual programming is the only working choice. Because the question is similar to any other software-engineering project. As an input we have a system specification, for example “program a software modul which simulates strategy aspects of a soccer play”, and the result is after some man years working code which fulfills the specification.

The concept of “Action model learning” is discussed sometimes in the literature, but it is only a future vision not a technology which is available today. The hope, that an action model can be derived from plan-traces without handwritten code will be disappointed.

1.8 Combining autonomous and remote-control

It is obvious, that writing a software-package which controls a robot is possible. Because, a robot is a machine which can be driven by commands, and a software can generate these commands. It is only a detail question, how exactly such a program will look like. In real robotics projects this is the bottleneck number one. That means, on a theoretical level it is often clear, what the robot should do, but the users are not good in programming or have not enough time to realize the software itself.

A possible answer is a semi-autonomous control system. The human operator takes first a joystick for controlling an underwater vehicle, and only in the second step the robot is following the demonstration. [11] The system described in the paper is at first hand a teleoperated underwater vehicle. That is according to the definition not a real robot, but it is more a remote controlled toy-boat. The advantage is, that writing the software for a human-in-the-loop simulation is relatively easy. The second part of the project (autonomous control) is postponed to the step #2. That means, even this subproject fail, it is always possible for a human to control the system manually.

The second advantage of this step wise programming is, that a concept like “Learning from demonstration” answers the question, how the Artificial Intelligence will work. The task can be specified under the term tracking. That means, the robot has the obligation to reproduce the human task. Implementing this task in software is not easy, but it is possible. In the above cited paper the authors are using the PDDL language for the high-level-planner and DMPs (parametrized dynamic movement primitive) for the lowlevel part.

But let us go into the details. The boat can be controlled manually. That means the human operator has a joystick and lots of keys he can press, and the boat has a certain task, for example “docking”. The process is the same like playing a computergame, that means the operator must press some keys, the robot reacts and the water in the tank increases the difficulty because the fluid produces disturbances. It needs usually more then one attempts until the human operator has mastered the task.

The strategy of the autonomous system can be sloppy described as Decision support system. That means, the AI is not really capable of controlling the robot, instead it is a support for the human operator. Sometimes the concept is called supervised autonomy, because the human operator is always in the loop and must observe what the robot is doing. How exactly works the system? It is mostly a controller. That is a piece of software which generates control-signal for the robot. The controller consists of a lowlevel and a high-level part. It is similar to a GOAP-planner which is used in computergames for controlling non-player-characters but is integrated in the overall robot-monitoring-system. The basic idea behind a controller is a planning task. That means, from the current situation a plan is generated to bring the system into a goal state. It is the same principle like a chess-engine works. There is a game-tree, different branches and an evaluation criteria. The difference is, that a robot-controller is more complicated then a chess-playing software. It contains more submoduls which are doing the hierarchical planning process.

Another interesting feature of mixed teleoperation is, that from the software-engineering side the project is relatively easy. That means, a teleoperated controlled robot is mostly a hardware problem which is understood very well. If hardware is available like a robotarm, a camera, an object and a joystick all the devices has to be connected and the system is ready. The joystick can be replaced by a data-glove and the manipulator by a larger model. That means, it is not necessary to write complicated control software for the task or have a theoretical understanding of robotics to pickup the ball. The main idea is to postpone the complicated task of writing the software to later state, namely for the Learning from demonstration. It is mostly the result of playing around with the system and try to improve their functionality a bit. Or to make the point clear: it is complicated to fail a telerobotics project. The reason why has to do, that a human-operator is capable of executing nearly all task. Remote controlling a machine is a robust way of interaction with the environment.

Suppose the manual teleoperation works, what is the next step? The next more advanced form is to utilize a pddl planner in the loop. A task model is formulated in the pddl language and this calculates some decision in realtime. Such a system is not a real robot system, because pddl is only a basic form of planning, but it is a good transition from a pure human-controlled system into a semi-autonomous system.

2 Example

2.1 HTN Planner

The difference between a Hierarchical task network and a behavior tree is not easy to understand. Both concepts are about Artificial Intelligence and Game-programming, but in general a HTN-planner is superior. Perhaps a simple example in sourcecode makes sense to grasp the idea.

The figure [fig:Textadventure] shows a compact C++ class which implements a textadventure. The user can send to the engine different commands like “init”, “open-door” and so on. After the command is parsed, the engine changes internal variables. In the concrete example, the user must first open the door until he can enter the room. The C++ class is not a behavior tree which says which commands must be executed in sequence, instead it is a game-engine who accepts commands and prints out the internal state. The user can play around with the game and send different commands to the engine in the hope, that he will reach the goal. It is not the HTN planner itself, but the domain model which can be used by a HTN-planner. The question is: which commands must be executed to bring the system into a goal-state.

class Textadventure {
public:
  std::string door;
  std::string position;
  void action(std::string name) {
    std::cout<
    if (name=="open-door") door="open";
    if (name=="close-door") door="close";
    if (name=="go-in") {
      if (door=="open") 
        position="inside-room";  
    }
    if (name=="go-out") {
      if (door=="open") 
        position="outside-room";  
    }
    if (name=="init") {
      door="close";
      position="outside-room";
    }
    if (name=="show") {
      std::cout<<"* ";
      std::cout<
      std::cout<<"\n";
    }
  }
  Textadventure() {
    action("init");
    action("show");
    action("go-in");
    action("show");
    action("open-door");
    action("show");
    action("go-in");
    action("show");
  }
};


Textadventure visual

The interesting aspect is the hierarchy of actions. On the lower-level the player can open and close the door and on the high-level layer he can enter and leave the room. Let us make a practical example. The game starts with a fresh instance and the agent want's to enter the room. He types in “action("go-in");” but it doesn't work, because the game-engine prevents that the player can direct manipulate his position. Instead there is a build-in game-mechanics, that means on the low-level the player must first open the door until he can execute the high-level “go-in” command.

The problem is not to solve the game automatically, this can be done by a brute-force sampler in under a second. The problem is to describe the domain in a machine-readable form. That means to program the game-engines on the lower-level and on the higher-level. A game-engine is a software-modul which specifies what will happen if the agent executes an action.

2.2 Car physics

Suppose we want to program a car parking controller, what is the best practice method? At first, we need a simulator with a top down physics engine. The idea is not to use a real car, but testing out the controller in a computer-game. A realistic physics engine is also helpful because it allows to detect collisions. But there is only one problem: a working simulator isn't equal to a working controller. A simulator means only, that we can control the car with a joystick, but the aim was to get an autonomous car.

Let us investigate what the problem is. If the car is visible in the simulator we can test out different plans, for example a trajectory for parking. The question is: how does look the trajectory for getting a certain goal? This is usually answered with a brute-force sampling planner, that means we are testing out 1 million trajectories and evaluate what will happen. So we can generate the parking trajectory like a chess engine works. There is only a minor problem. The CPU-consumption would be very high. Even if we are taken a modern efficient physics engine like Box2d or bullet, a normal computer is not able to evaluate more then 100 trajectories in a second. That means, it is not possible for testing out 100 million trajectories, it would take years for doing so.

Car-parking-scene


The answer to the problem is called “hierarchical task network”. It is a planning technique which constructs a layered simulator and planning on different levels. Let us make an example. Figure [fig:Car-parking-scene] contains three scenes of a parking maneuver. The standard programming technique to implement the game is a realistic physics engine, because it is highly accurate. The alternative is to program a simplified physics engine from scratch. This engine doesn't have a collision check or sophisticated force calculation. Instead it works on a symbolic level.

The game starts with the init-keyframe. That means, there is a car and a parking lot. The user has different commands he can enter. The command “u-turn” does not activate a complex AI-controller which is doing the u-turn maneuver with a detailed trajectory. No, “u-turn” means only to change the direction of the car direct. That means, in the game the position of the lamps of the car will be changed to the new position, that is all. If the user enters “u-turn” the game-engine simply flips the car. The next command the user can enter is “parking”. Like in the example before, it is a very basic command, if the user enters “parking” the position of the car is simply changed to the parking lot. That physics engine works in a reduced form, it is similar to a pddl specification.

As a consequence we have a car-parking game, but a very abstract kind of game. Playing around with the game isn't generating the detailed trajectory, but only the subgoals. It is a way of storing knowledge and to enable high-level-planning. Let us now investigate how “parking” works. The user enters two commands: “u-turn”, “parking”. That's all. With the first command he flips his car, and with the second command he moves the car into the lot.

Until now the idea may look a bit useless, because we need a detailed motion controller and not a subgoal generator. But the described abstract game is an important step into this direction. We can use the high-level-game for constructing the low level controller. If we know, what the subgoal is, we can do the planning on the realistic physics engine. The question is not longer: what is the overall trajectory, the question is only “how to realize a u-turn”? That means, the high-level commands are equal to skill primitive. Calculating the correct trajectory for reaching out these skills is easier then planning the complete domain. It is possible to use a realistic physics engine like Box2D for calculating such subgoals.

References

[1] Luciana Benotti, "DRINK ME: Handling actions through planning in a text game adventure", XI ESSLLI Student Session  (2006), pp. 160--172. https://cs.famaf.unc.edu.ar/~luciana/content/papers/files/benotti06.pdf

[2] Donna Byron, Alexander Koller, Kristina Striegnitz, Justine Cassell, Robert Dale, Johanna Moore, and Jon Oberlander, "Report on the first NLG challenge on generating instructions in virtual environments (GIVE)", in Proceedings of the 12th european workshop on natural language generation (, 2009), pp. 165--173. http://www.aclweb.org/anthology/W09-0628

[3] Aaron St Clair, Carl Saldanha, Adrian Boteanu, and Sonia Chernova, "Interactive hierarchical task learning via crowdsourcing for robot adaptability", in Refereed workshop Planning for Human-Robot Interaction: Shared Autonomy and Collaborative Robotics at Robotics: Science and Sys… (, 2016). https://people.csail.mit.edu/cdarpino/RSS2016WorkshopHRcolla/abstracts/RSS16WS_17_InteractiveHierarchicalTask.pdf

[4] Richard E Fikes and Nils J Nilsson, "STRIPS: A new approach to the application of theorem proving to problem solving", Artificial intelligence  2, 3-4 (1971), pp. 189--208. http://ai.stanford.edu/~nilsson/OnlinePubs-Nils/PublishedPapers/strips.pdf

[5] Andrew Gargett, Konstantina Garoufi, Alexander Koller, and Kristina Striegnitz, "The GIVE-2 Corpus of Giving Instructions in Virtual Environments.", in LREC (, 2010). http://cs.union.edu/~striegnk/papers/striegnitz_conference_lrec_2010.pdf

[6] Andrew Garland and Neal Lesh, "Learning hierarchical task models by demonstration", Mitsubishi Electric Research Laboratory (MERL), USA--(January 2002)  (2003). http://www.merl.com/publications/docs/TR2002-04.pdf

[7] Alexander Koller and Ronald Petrick, "Experiences with planning for natural language generation", Computational Intelligence  27, 1 (2011), pp. 23--40. http://www.coli.uni-saarland.de/~koller/papers/ci-crisp-11.pdf

[8] Miguel Lozano, Steven J Mead, Marc Cavazza, and Fred Charles, "Search-based planning for character animation", in 2nd International Conference on Application and Development of Computer Games (, 2003). https://pdfs.semanticscholar.org/cba3/a6527f7a4e7209b463cc70bc300cda01b171.pdf

[9] Dipendra K Misra, Jaeyong Sung, Kevin Lee, and Ashutosh Saxena, "Tell me dave: Contextsensitive grounding of natural language to mobile manipulation instructions", in in RSS (, 2014). https://www.cs.stanford.edu/people/asaxena/papers/misra_sung_saxena_rss14_tellmedave.pdf

[10] Oliver Obst, Anita Maas, and Joschka Boedecker, "HTN planning for flexible coordination of multiagent team behavior", Fachberichte Informatik  (2005), pp. 3--2005. ftp://ftp.uni-koblenz.de/pub/outgoing/Reports/RR-3-2005.pdf

[11] Narćıs Palomeras, Arnau Carrera, Natàlia Hurtós, George C Karras, Charalampos P Bechlioulis, Michael Cashmore, Daniele Magazze…, "Toward persistent autonomous intervention in a subsea panel", Autonomous Robots  40, 7 (2016), pp. 1279--1306. https://www.researchgate.net/profile/Charalampos_Bechlioulis/publication/283339628_Toward_persistent_autonomous_intervention_in_a_subsea_panel/links/587574d108ae8fce492823bc/Toward-persistent-autonomous-intervention-in-a-subsea-panel.pdf

[12] RM Simpson and W Zhao, "Gipo graphical interface for planning with objects", International Competition on Knowledge Engineering for Planning and Scheduling  (2005), pp. 34--41. https://pdfs.semanticscholar.org/5ef6/d24c8f496963a40b4939b5f3e012c977edca.pdf

[13] Qingxiaoyang Zhu, Vittorio Perera, Mirko Wächter, Tamim Asfour, and Manuela Veloso, "Autonomous narration of humanoid robot kitchen task experience", in Humanoid Robotics (Humanoids), 2017 IEEE-RAS 17th International Conference on (, 2017), pp. 390--397. http://h2t.anthropomatik.kit.edu/pdf/Zhu2017.pdf