Der simple Grund warum bis ca. 2010 es unmöglich war Roboter zu bauen und zu programmieren liegt darin, dass unklar blieb wie genau komplexe Aufgaben in Software algorithmisch bewschrieben werden sollten. Wenn z.b. ein Roboter den kürzesten Weg zum Ziel finden soll um dort einen Gegenstand präzise zu greifen benötigt der Roboter eine extrem komplexe Steuerungssoftware aus vielen tausend Zeilen von Code. Dieser Code muss irgendwer programmieren.
Soll der Roboter komplexere Aufgaben lösen z.b. bei selbstfahrenden Autos oder beim biped walking erhöht sich der Programmieraufwand weiter. Zwar konnten KI Forscher bis 2010 grundsätzlich den benötigten C/C++ Softwarestack programmieren udn mögliche Fehler darin korrigieren, nur eben in einem sehr langsamen Tempo mit den üblichen 10 lines of code pro Tag für neu zu erstellende Software. Es verwundert wenig das größere Robotik-Projekte mit diesem Entwicklungstempo unmöglich zu realisieren waren.
Die Antwort auf das Dilemma besteht keineswegs darin über genetische Algorithmen die Software sich selbst verbessern zu lassen sondern die benötigte Technologie um Roboter zu programmieren besteht in der Externalisierung von Intelligenz. Je mehr Aufgaben der Roboter an eine Umwelt deligiert desto schlanker ist die verbleibende Robotik-Software. Ein anderer Begriff für Externalisierung lautet Teleoperation. Das heißt die Wegfindung oder die Greifplanung eines Roboters wird nicht länger in Software realisiert sondern durch einen menschjlichen Bediener außerhalb des Roboters.
Anstatt das Für und Wieder von Teleoperation zu untersuchen ist es wichtiger sich die Sprachinterfaces zu kümmern die zwischen dem Roboter und dem menschlichen Bediener benötigt werden. Dieses Interface funktioniert in bidirektionaler Richtung: der Roboter reagiert auf gesprochene Kommandos und der Roboter sendet verbale Statusinformation an den Menschen zurück.
Mit Hilfe eines natürlichsprachlichen Interfaces lassen sich einerseits hochkomplexe Roboter realisieren wie z.b. selbstfahrende Autos, während gleichzeitig die benötigten Codezeilen gering sind. Die Aufgabe der Robotersoftware ist weniger das Steuern des Roboters, sondern das Ziel ist ein Sprachinterface bereitzustellen damit der Roboter effektiv aus der Ferne gesteuert werden kann.
Bis ungefähr zum Jahr 2010 erschien das Konzept von ferngesteuerten Robotern absurd. Es wurde nichtmal kritisch distanziert von der KI Community diskutiert sondern überhaupt nicht thematisiert. Der Konsens stattdessen lautete3 autonome Roboter zu bauen die ohne Teleoperation funktionieren. 100% aller Robotikwettbewerbe wie Robocup, micromouse, Carolo-Cup oder die Darpa Grand challenge arbeiteten nach diesem Prinzip.
August 09, 2026
Grenzen autonomer Robotik
August 08, 2026
Sprachspiele im Kontext von Robotik
Die Erforschung künstlicher Intelligenz drehte sich lange Zeit um die Frage welche Art von Softwareframework, Algorithmus oder Programmiersprache benötigt wird um denkende Maschinen zu realisieren. Die Annahme hinter dem Cyc Projekt von Douglas Lenat lautete, dass eine Ontologie der Grundbaustein sei, das Cam-Brain Projekt von Hugo de Garis unterstellte dass im Kern ein neuronales Netz benötigt wird während Edward Feigenbaum vermutete dass sich Künstliche Intelligenz mit Hilfe von Expertensystemen realisieren läst.
Im Laufe der Zeit wurden sehr viele gegensätzliche Technologien und Annahmen entwickelt um Künstliche Intelligenz zu verwirklichen. Ein Sprachspiel im Sinne von Ludwig Wittgenstein ist nur ein weiterer Vorschlag unter vielen. Dennoch lohnt es sich, das Thema nähter zu untersuchen. Weil das Prinzip von Sprachspielen deutlich anders ist als z.b. ein Expertensystem oder ein neuronales Netz.
Sprachspiele sind ein formalisiertes Interface zwischen interner und externer Realität von Systemen. Es ist ein Test, ob die Spielteilnehmer Sprache erzeugen und verstehen mit derer man die externe Realität beschreibt. Ein typisches Beispiel ist das "Ich sehe was du nicht siehst Spiel". Dabei beschreibt Person A ein Objekt aus der Realität anhand von Eigenschaften, er formuliert aussagen wie "das Objekt ist rund, steht in der Küche, hat Beine" und daraus folgert Person B "es ist ein Tisch".
Die Gemeinsamkeit von allen Sprachspielen ist, dass Sprache in einem interkationsspiel genutzt wird um auf die Wirklichkeit zu referenzieren. Es geht immer darum, Dinge abzufragen die in der Realität vorkommen, Anweisungen geben was in der Realität zu tun ist oder sonstwie Sprache und Wirklichkeit in Beziehung zu setzen. Sprachspiele sind eine gute Möglichkeit eine Fremdsprache zu erlernen und dienen dazu den Wissensstand zu erfassen. Wenn eine Person oder ein Computer in einem Sprachspiel eine hohe Punktzahl erreicht, ist diese Person mit der jeweiligen Sprache vertraut.
Anders als die eingangs erwähnten Technologie zur Realisierung von künstlicher Intelligenz wie Onotologien oder Expertensysteme sind Sprachspiele nicht technisch definiert sondern haben ihren Ursprung in der Philosophie und der Linguistik. Es geht um Themen wie interaktivität, Sprache, Wirklichkeit. Die Annahme lautet dass innerhalb dieser Begriffe Künstliche Intelligenz möglich ist, das also denkende Maschinen ein Interface zwischen internem System und externer Realität sind.
Seit den 1980er wurden Brettspiele als Testumgebung für denkende Maschinen genutzt. Schach ist ein sehr altes Beispiel für das Künstliche Intelligenz realisiert wurde, aber auch andere Spiele wie Tic Tac Toe, Dame, Backgammon und Go sind klassische Umgebung zur Erforschung von KI Algorithmen. Leider haben diese Brettspiele den Nachteil dass sie nicht gut nach oben skalieren. Ein Computer der perfekt Schach spielt ist nicht automatisch im Stande einen Roboter zu steuern. Deshalb eignen sich die erwähnten Brettspiele nur sehr eingeschränkt dazu KI näher zu erforschen.
Sprachspiele kann man als neuartiges Gesellschaftsspiel verstehen was ähnlich wie Schach Regeln folgt aber viel besser nach oben skaliert. Ein Computer der das "Guess what" Sprachspiel beherscht ist zugleich auch in der Lage mit diesem Wissen einen Roboter zu steuern. Scheinbar können Sprachspiele die Kernidee von Künstlicher Intelligenz viel besser formalisieren als frühere Brettspiele.
August 07, 2026
Old school artificial intelligence
Before the advent of large language models, there was lots of AI related research available which was mostly ignored by the public. Typical topics in robotics from the past was NEAT neuroevolution, model predictive control and genetic algorithms. All these concepts were working with computational paradigm. At first the fitness function was defined, e.g. the robot gets +1 reward for moving on step forward, and then an algorithm modifies the parameters of the neural network. In case of model predictive control it was even possible to plan one step ahead, so that the robot was able to use a trial&error strategy.
These concepts can be called old school artificial intelligence because they were working with computers and algorithms as the core element. The assumption was that artificial intelligence is equal to search, so its a mathematical optimization problem. Instead of finding the prime numbers from 1 to 1 million and find 1000 digits of the PI constant, the goal was to use raw processing power to solve robotics problems.
The mentioned NEAT and genetic algorithms strategy were the core principle of AI in the past. Large scale projects like CAM brain and the fifth generation computers in Japan were built around these subjects. The researchers have treated AI problems as mathematical optimization problems by transforming robot motion planning into a mathematical state space first, and then use algorithms to solve these problems.
What was unclear until around 2010 was, that all these strategies have failed. Mathematics and algorithms in general are not powerful to solve robotics problems. This insight is even today some sort of blasphemy because mathematics and algorithms are the building blocks in computer science. If both concepts are rejected as weak then the amount of remaining strategies is small and even empty.
The major result of old school artificial intelligence was not to build robots but to proof that mathematical optimization is a dead end. It can be formulated from a theoretical perspective, that all interesting robotics problems are np hard, and that genetic algorithms including NEAT are not able to solve this problem category. Therefor AI can't be realized.
Its possible to replicate the failure of NEAT and similar approaches with today's software libraries. At first a robotics problem gets formulated, e.g. a walking robot in the OpenAI gym which is a 2d simulator. then a state of the art neuroevolution library is utilized to control the robot. A handcrafted fitness function will reward the algorithm so that the optimization can evolve into improved biped walking generator. Unfortunately, the described setup won't work in reality. Even if the robot walks some steps forward the robot will struggle on larger obstacles. At the same time, the CPU consumption is very high.
At first the fitness function of the robot will improve, but then there is no further improvement available and its impossible to fix the problem. Some reinforcement learning advocates are explaining that with parameter tuning the problem can be solved, other NEAT users claim that with future algorithm update the issue might be solved. But the sad reality is, that mathematical optimization in general fails for robotics problem. No matter if NEAT, genetic algorithm or Q learning was utilized, none of these strategies can control a biped walking robot.
Such kind of insight isn't the end of robotics research but it provides meaningful facts about the limitation of old school algorithms. It shows on a practical example, what sort of AI algorithm isn't working in reality. This opens the path for more advanced AI techniques beyond classical computer science.
August 03, 2026
KI als Berechnungsprozess
Zumindest früher und in einem universitären Kontext wurde Künstliche Intelligenz erstaunlich präzise definiert. Rückblickend betrachtet war diese Definition eine Sackgasse welche paradoxerweise das Entstehen von leistungsfähiger Robotik sowie von Large language modelle verhindert hat anstatt zu ermöglichen. Dennoch ist es wichtig sich mit dieser philosophischen Sackgasse näher auseinanderzusetzen.
In einem älteren Foliensatz aus dem Jahr 2005 wurde erläutert dass Künstliche Intelligenz ein Berechnungsprozess sei. [1] page 3. Um ein Problem zu lösen müsse man im Problemraum suchen. [1] page 5
Es werden mehrere Beispiele genannt wie das "Affe und Banana Problem" und "Der Turm von Hanoi" und es werden Lösungsalgorithmen vorgestellt wie uniformierte Suche, Backtracking und Hill climbing mit Bewertungsfunktion.[1] page 12
Insgesamt ist die Darstellung richtig und für die damalige Zeit state of the art. Foliensätze von anderen Universitäten waren ähnlich aufgebaut. Die Grundannahme lautet dass ein klar definiertes Problem existiert was sich über Suchalgorithmen lösen lässt. Als bestes Suchverfahren gilt Hill climbing mit Bewertungsfunktion was auch als heuristische Suche bezeichnet wird.
Die Aufgabe des Computerprograms besteht darin diese Suche auszuführen. Also eine Berechnung durchzuführen welche als Algorithmus definiert wurde. Demnach ist KI die Anwendung einer CPU um Probleme zu lösen.
Was zumindest in der Vergangenheit nicht bekannt war oder verdrängt wurde ist, dass das Verfahren nur bei sehr simplen Problemen funktionoiert wie Finden eines Weges in einem Labyrinth oder Computerschach. Sobald die Probleme geringfügig anspruchsvoller werden, z.B. sehr große Labyrinthe, komplexere Brettspiele oder gar Robotik versagt das Verfahren. Ein typisches Beispiel ist Greifplanung für eine Roboterhand. Selbst wenn eine Bewertungsfunktioni existiert gelingt es nicht mittels Suche im Problemraum eine Lösung zu ermitteln. Der Suchraum ist viel zu groß.
[1] Joachim Funke: Seminar: Denken und Problemlösen, Ruprecht-Karls-Universität Heidelberg, 2005, https://www.psychologie.uni-heidelberg.de/ae/allg/lehre/050217_Computermodelle.pdf
August 01, 2026
KI entwickeln ohne die Verwendung Künstlicher Intelligenz
Während der 1980er Jahre versuchte die KI Forschung Computern das Denken beizubringen. Vorhandene Programmiersprachen wie C oder Pascal wurden als nicht leistungsfähig genug betrachtet und durch dezidierte 5th generation languages wie Prolog und Expertensystems shells ersetzt. Die Zielstellung war, vorhandenes Wissen leichter auf einen Computer zu übetragen damit anschließend der Computer das Problem zu lösen vermag.
Einen ähnlichen Ansatz verfolgt das Semantic Web. Auch dort wurde die komplexe Wirklichkeit in Fakten gespeichert und dann von Rule-Engines ausgewertet. Man kann also semantic web als die logische Weiterentwicklung einer Expertensystem shell betrachten.
Im wesentlichen ging es darum, Teile der Wirklichkeit in ein Computerprogramm zu konvertieren. In diesem Computerprogramm konnten dann logische Schlüsse gezogen werden, eine Suche durchgeführt werden und Entscheidungen getroffen werden. Die Sprache Prolog ist sehr gut dafür geeignet, ein logisches Wissensnetzwerk zu erzeugen. Man gibt zuerst die Fakten in das Program ein und kann dann fehlende Informationen daraus folgern.
Die Stillschweigende Annahme hinter diesen Technologien bestand darin, dass eine KÜnstliche Intelligenz zwingend auf einem Computer ausgeführt wird in Form einer Software. Beim Semantic web gab es eine Wissensbasis gespeichert im Computer plus eine Rule engine, bei der Prolog Sprache gab es die Software einerseits und den Prolog interpreter andererseits. Diese strukuturelle Kopplung aus Künstlicher Intelligenz und Informatik ist historsich gewachsen, nur leider war sie zugleich für den KI Winter Anfang der 1990er Jahre verantwortlich.
Solange man die Zielstellung verfolgt, ein Computer programm zu erstellen oder eine Wissnsbasis anzulegen die von einem Computer program ausgewertet wird entsteht ein möglicher Reality gap. Also der Unterschied zwischen der Wirklichkeit und dem Computermodell was diese Wirklichkeit simuliert. Die einzige Methode diesen gap zu minimieren besteht darin, die Anzahl der regeln im Prolog Program zu erhöhen. Dadurch steigt die Komplexität der Software.
Die Antwort zum Reality gap problem lautet: Fernsteuerung. Es gibt also keine Robotik-Software welche den Roboter steuert sondern es gibt einen menshclichen Bediener. Und es gibt auch kein Expertsystem was eine Diagnose erstellt sondern ein menschlicher Experte erstellt die Diagnose. Durch die simple Forderung nach Fernsteuerung löst man alle Probleme der Künstliche Intelligenz der 1980er. Man tauscht das vorherige Berechenbarkeitsparadigma aus zugunsten eines Kommunikationsparadigmas. Der Computer ist nicht mehr dafür zuständig Programcode auszuführen, sondern der Computer übertragt Funksignale von A nach B.
Diese veränderte Rollenzuweisung entspricht weitaus besser der technischen Leistungsfähigkeit. Wenn die einzige Aufgabe des Computers darin besteht Signale weiterzuleiten sind die benötigte Hardware- Software- und Algorithmen trivial. Gleichzeitig entsteht ein neues Dilemma: Fernsteuerung bedeutet, dass kein Computer die entscheidung trifft sondern ein Mensch. Das wird üblicherweise nicht als KI betrachtet. Sobald ein Mensch mit Fernsteuerung den Roboter steuert, hat man eben keinen autonomen Roboter mehr der über eine Software Entscheidungen trifft sondern man hat ein normales RC Car realisiert.
Erst mit einem erweitereten KI Begriff der distributed cognition mit einschließt wird ein ferngesteuertes Auto zu einer Künstlichen Intelligenz umdefiniert. Distributed cognition ist eine relativ neue philosophische Richtung die davon ausgeht, dass Intelligenz sich aus der Kommunikation der Invididuum ergibt. Demnzufolge kann eine einzelne CPU in einem isolierten Roboter nicht intelligent sein, erst wenn diese CPU mit einer anderen CPU oder einem Menschen interagiert, entsteht Intelligenz.
July 31, 2026
KI Definition in der Vergangenheit
Eine häufig formulierte Definition von Künstlicher Intelligenz früher lautet, dass es darum geht menschliches Denken mit einem Computer zu simulieren. Diese Definition ist abgeleitet vom Turing Test, kann aber auch ohne einen mathematischen Hintergrund formuliert werden. Auch außerhalb der Informatik ist ein Roboter eine Maschine die so aussieht wie Mensch und das selbe leistet. Es geht also darum das menschliche Denken in einem Computer nachzubauen.
Als Einstieg mag diese Definition nützlich sein, allerdings gibt es das Problem dass unklar bleibt wie genau ein Computer einen Menschen nachbilden soll. Computer werden über Software programmiert und diese Software muss sehr präzise Angaben enthalten was der Computer tun soll. Genau an dieser präzisen Vorstellung bezüglich menschlichen Denkens fehlt es. Es gibt zwar durchaus Ansätze den Menschen zu untersuchen mit Hilfe von Biologie, Psychologie und Soziologie nur lassen sich die dabei gewonnenen Erkenntnisse nicht in Software übertragen. Als Folge wird Künstliche Intelligenz zu einer theoretisch philosophischen Betrachtung verkürzt und kann nicht durch Experimente falsifiziert werden.
Die Definition künstlicher Intelligenz bildet den Rahmen innerhalb derer die Erforschung vorangetrieben wird. Die obige Definition "KI=Menschen nachbauen" ist zu ungenau. Dadurch wird KI verunmöglicht.
Hier ist eine modernere Definition von KI, welche den Fokus auf Interaktion mit der Umwelt richtet:
"Künstliche Intelligenz ist die Fähigkeit eines Systems, in Echtzeit mit seiner Umgebung (einschließlich Menschen) zu interagieren, indem es sensorische Eingaben verarbeitet, Handlungen plant und ausführt sowie durch Feedback (z. B. Sprache, Gesten oder andere Signale) sein Verhalten anpasst, um zielgerichtete Aufgaben zu erfüllen oder Kommunikation zu ermöglichen."
July 29, 2026
KI als Mapping problem in der DIKW pyramide
Zumindest bis in die 2000er Jahre wurde Künstliche Intelligenz so definiert, dass menschliches Denken mit Hilfe eines Computers simuliert werden soll. Aufbauend auf dieser Definition wurden mehrere Strategien und Algorithmen diskutiert, die jedoch das selbst definierte Ziel nicht zu erreichen vermochten. Der wohl erfolgreichste Algorithmus der KI Forschung bis zum Jahr 2000 war der bekannte Minimax Algorithmus um Schach von einem Computer spielen zu können. Dieser Algorithmus war so erfolgreich, dass er sogar Großmeister schlagen konnte. Nur, Minimax lässt sich nicht auf Robotik-Probleme verallgemeinern.
Eine neuere und weniger bekannte Definition von Künstlicher Intelligenz lautet dass es ein Zuordnungsproblem ist zwischen low level daten und high leven natürlicher Sprache. Dieses Zuordnungsproblem entsteht innerhalb der DIKW pyramide (data, information, knowledge, wisdom). Vielleicht ein kleines Beispiel:
Über Motion capture wird die Bewegung eines menschlichen Aktors aufgezeichnet. Die x/y/z position der Mocap Marker sind Daten und gehören zum Data layer der DIKW pyramide. Jetzt soll über Mustererkennungsverfahren ermittelt werden, welche Pose der menschliche Aktor gerade ausführt. Diesre kann gehen, rennen, sitzen oder springen. Diese Labels werden im Information layer der DIKW pyramide gespeichert. Aufgabe für die Künstliche Intelligenz ist es jetzt zwischen beiden layern eine Zuordnung herzustellen. Es geht also um die Frage wie eine mathematische Realität auf eine linguistische Realität projiziert werden kann.
Die naheliegende Frage lautet: Warum muss der mathematische Raum der Mocap Marker auf einen linguistischen Raum der textuellen Label projiziert werden? Kann man die Steuerung von Robotern nicht auch eleganter/einfacher beschreiben?
Die Notwendigkeit der Sprachlichen Enkodierung ergibt sich aus dem Kommunikationserfordernis. Es reicht nicht, die mocap Marker als x/y/z position zu speichern, sondern zusätzlich soll einem außenstehender Beobachter in natürlicher Sprache übermittelt werden, was genau der Aktor gerade tut. Ziel ist weniger ein System intern zu beschreiben, sondern Zielstellung ist mit einem 2. System zu kommunizieren.
KI Forschung bis ca. zum Jahr 2000 war geprägt von der closed system Hypothese. Ziel war stets das interne Funktionieren eines Computers zu optimieren durch bessere Algorithmen, schnellere Programmiersprachen oder spezielle Datenbanken wie Cyc. Im Gegensatz dazu verfolgt KI Forschung ab dem Jahr 2000 das Ziel, die Kommunikation zwischen zwei Systemen zu verbessern. Die Frage lautet stets: wie kommunizieren zwei Menschen miteinander? Wie interagiert ein Mensch mit einem Roboter? Welche Interaktionsspiele zwischen einem Speaker und einem Hearer gibt es?
Selbstverständlich ist die obige Beschreibung stark vereinfachend. Auch vor dem Jahr 2000 gab es Versuche die Mensch-Maschine Interaktion zu verbessern:
- 1968,SHRDLU natural language understanding by Terry Winograd
- 1987,Vitra visual translator
- 1998,Rocco Robocup commentator by Dirk Voelz
Man könnte z.B. SHRDLU als Vorläufer heutiger Vision Language action (VLA) Modelle beschreiben. Anzahlmäßig waren jedoch KI Projekte, die natürliche Sprache als Kommunikationsschnittstelle verwendeten eher die Ausnahme. Es war vor dem Jahr 2000 unklar, dass dies wichtig ist und es war unklar wie man eine solche Schnittstelle technisch realisieren könnte. Hinzu kommt dass ein interdisziplinärer Ansatz bestehend aus Mathematik und Lingustik vor dem Jahr 2000 unüblich war.
July 28, 2026
Graph traversal with a head up display
The perhaps most simple example for a head up display is a graph traversal problem of a robot. The robot moves inside a graph and should reach a target node.
The AI for the robot works with a head up display. There is a text box at the bottom showing the inner voice of the robot. The inner voice determines at which position the robot is, which nodes are in the near, what the target node is, and which action should be taken next.
A mathematical problem, graph traversal, gets converted into a textual problem. Textual means, that the head up display is using words to describe the reality. possible words are [currentnode, goalnode, nextnode, distance_to_goal]. These words and events are used to describe the game state from a high level perspective. The text box ensures that the inner voice was implemented correctly. That means, the AI isn't solving an optimization problem and its not running an algorithm, but the main task for the AI is to generate textual output in the head up display and talk to the human operator.
July 27, 2026
The paradigm shift in robotics around the year 2010
According to published research papers, the year 2010 was a turning point in robotics research. After this year, higher effort was put into human to robot interaction with natural language. Projects from this time span were:
- M.I.T. forklift by Stefanie Tellex
- Marco route instruction following by Matt MacMahon
- Word2vec algorithm by Tomas Mikolov
These projects were started from around 2008 until 2013. Many smaller projects also tried to use natural language for robot control.
This development was different from robotics research until 2010. The years before this year, there was a search for sophisticated algorithms available like training algorithms for neural networks, path planning algorithms and SLAM algorithms. The search for novel algorithms was working with the same principle how common computer science is working. The idea was, that AI gets implemented on a computer, computers need an algorithm and the consequence is to develop dedicated robotics algorithms for solving tasks.
The problem with the algorithm centric perspective until 2010 was, that all these developed techniques were not powerful enough. Even advanced probabilistic path planning algorithms implemented on a multi core CPU are not able to control a warehouse robot. The problem is the reality gap. The robot assumes a different reality than the real reality and the algorithm can't bridge the gap. It makes no sense to program more additional software modules or create a larger database for the robot, but the principle of autonomous robotics in general has to be questioned.
This paradigm shift took place in published academic literature around the year 2010. Research papers written after this date put a higher emphasis on teleoperation and grounded language for robot control. The idea is, that the source of wisdom is located outside of the robot as a human operator and the task is to get access to this knowledge by asking the human in natural language.
Even if the principle sounds inaccurate, it can be scaled up towards more complex scenarios. In the easiest case the human to robot interface works with a list of predefined commands, in a more advanced setup a neural network can parse the instructions. In contrast to figure out algorithms for autonomous robots, the new paradigm is to build language parser and see a robot as an open system.
The paradigm shift around the year 2010 allows to create artificial intelligence. Nearly all the former problems in robot control can be solved with the open system paradigm. Its only a detail problem how to create a high level user interface, so that the human operator can provide general statements like "tidy up the kitchen" and the robot is doing the full task by its own.
Roughly spoken, robotics programmed after the year 2010 are entirely teleoperated. There is a human operator in the background who gives instructions, or the former human operator instruction list was translated into a computer program who talks with the robot as large language model. in all the cases the robot interacts with a higher instance outiside of the robot and the artificial intelligence is located in the language interface.
July 25, 2026
Engine for grounded language
One possible explanation why the symbol grounding problem has emerged late in the history of computer science is because the theory is difficult to realize on a computer. Suppose natural language is important for robot control, the problem is create a language parser which is working for a concrete domain.
A possible command for a robot might be "Move until obstacle and then stop". Each of the words is stored as a string, but it remains unclear who to process the instruction into actions for a robot. The reason is, that the sentence is formulated in English but computers need a programming language as input. Even if every word is encoded as a number, it doesn't make sense to submit an array with numbers to the robot because its not possible to add or subtract the values in a meaningful way.
In general the problem is how to convert natural language into a computer program. Without solving this issue, the symbol grounding problem remains only a philosophical problem without any practical consequences.
The good news is, that the problem of programming a parser can be solved. Not with tools from computer science but by using techniques from linguistics, namely language games. Instead of treating language parsing as an algorithm problem, the idea is to invent around words a puzzle game. Typical language games are:
- Name guessing game. Player1 points to an object in the reality, and Player2 has to tell the name
- NPC quest game, a non player character in a role playing game formulates a quest like "bring me the sword from the wood" and the player has to fulfill the task
- bounding box game, player1 says a word like "table" and player2 has to draw a bounding box around this object
All these language games are located outside of computer science. They have nothing to do with algorithms, programming language nor existing robotics libraries, but they are games played with 2 human players.
The interesting situation is, that its possible to simulate the games with a computer. The software encodes the rules of the language game, determines the score for the human player and then the player can take action inside this game.
The problem is not how to program a certain parser, but the problem is how to formulate the game outside of a computer first. A well formulated game can be implemented in a software with ease. The programmer needs only the specification of the game including its rule, and then its possible to program the game with python. The only requirement is, that the computer works like the original language game. The programming workflow is identical to implement card games and board games on a computer.
July 24, 2026
Die späte Entdeckung der Schrift im Kontext von Robotik
Die Schrift ist eine sehr alte Erfindung der Menschheit. Die erste Bilderschrift, die Ägyptischen Hieroglyphen entstanden um 3200 v. Chr. Insgesamt ist Schrift und natürliche Sprache sehr detailiert erforscht. Es gibt umfassende Wörterbücher, Darstellungen welche die Geschichte der Sprache zeigen und Untersuchungen bezüglich Wortherkunft.
Vereinfacht gesagt sind Wörter Referenzsysteme zur Realität. Substantive stehen für Objekte wie "Himmel, Tisch, Apfel", Adjektive stehen für Tätigkeiten wie "Laufen, springen, geben" und Adjektive werden als Eigenschaftswörter verwendet wie "gelb, groß, schnell, feucht". Das Wissen bezüglich Wortarten und die Nennung von Beispielwörtern ist banal, allerdings nur für die Sprachwissenschaft selber. Im Bereich Computerwissenschaft und Mathematik wurde natürliche Sprache lange Zeit ignoriert. Es gab zwar Versuche chatbots zu programmieren, aber das war nur ein Teilbereich der Künstlichen Intelligenz.
Erst in jüngerer Zeit stellte sich heraus, dass natürliche Sprache womöglich das fehlende Puzzleteil darstellt mit der man künstliche Intelligenz inbesamt realisieren kann. Und zwar indem man Sprache als Technologie verwendet. Insbesondere dessen Eigenschaft auf die Realität zu verweisen macht es zum idealen Abstraktionsmechanismus. Es müssen keine neuen Sprachen erfunden werden sondern vorhandene Sprachen wie English, Deutsch usw. bieten bereits ein umfassendes Vokabular was von Robotern ähnlich wie Menschen verwendet werden kann. Alles was eine Maschine dafür benötigt ist eine Übersetzungstabelle von Bildern zu Sprache und in umgekehrter Richtung. Mit Hilfe dieser Bild zu Wort Tabelle kann man einem Roboter ein Kommando geben wie "Fahre zum Tisch". Und der Roboter übersetzt den Satz dann in Bilder und in Aktionen.
Bis ungefähr zum Jahr 2000 hat die KI Forschung nach Algorithmen gesucht mit deren Hilfe sich denkende Maschinen konstruieren lassen. Typische Algorithmen waren Lernverfahren für neuronale Netze, SLAM Algorithmen zur Selbstlokalisierung, A* Pfadplanungsalgorithmen oder Momdel predictive control Algorithmen. Die Annahme lautete jeweils dass mit diesen Algorithmen denkende Maschinen konstruiert werden könnten. Diese Annahme ist jedoch falsch. Es liegt zusätzlich der Verdacht nahe, dass es generell keine Algorithmen gibt, die Künstliche Intelligenz erzeugen, weil ein Algorithmus per se nicht mächtig genug ist um Roboter zu steuern. Was man stattdessen verwenden könnte wäre natürliche Sprache als zentrales Koordinierungsinstrument. Sprache ist ein Interface zwischen den Wortsymbolen einerseits und der Realität andererseits. Dadurch kann die komplexe Realität in wenige Wörter komprimiert werden.
Erst durch diese Realitätskompression ist es möglich den Handlungsraum für Roboter zu verkleinern. Der Roboter plant nicht länger in einem 3d Raum mit Trajektorien sondern der Roboter plant mit Hilfe von Verben und Substantiven in einem abstrakten Sprachraum. Seit der Erfindung von Word embeddings wie Word2vec mag dieser Ansatz selbstverständlich klingen aber bis zum Jahr 2000 war der Fokus auf Sprache eine Revolution.
Noch immer steht natürliche Sprache ein wenig außerhalb der klassischen Computerwissenschaft. Es hat nichts zu tun mit Elektrotechnik, Mathematik oder Algorithmen, sondern die Sprachwissenschaft hat ihren Ursprung in den Geisteswissenschaften. Nicht nur in der Dewey Dezimalklassifikation wie sie in Bibliotheken verwendet wird, sondern auch in der Gliederung von Universitäten sind Geistes- und Naturwissenschaften unversöhnliche Gegensätze die getrennt betrachtet werden.
Open systems for robotics
Robotics in the past was organized with a closed system paradigm. A robot was described as a machine which consists of hardware, software and algorithms and the task for the programmer was to improve the internal mechanism of the robot. It was ignored that robots are communicating with the outside world. For example a robot might receives commmands by teleoperation and submits a status code to the operator. Such kind of interaction was mostly described as wrong path towards robotics because such a machine isn't autonomous anymore. There decision making isn't determined by the internal algorithm but from the outside which was seen as anti pattern in Artificial intelligence.
It takes decades until computer science has questioned the self created bias. Modern robotics is working as open system which means, that the robot gets information from sensors and from remote control. Also the robot interacts with human operators and is able to answer questions like "What object is visible in the camera?".
The transition from closed to open systems in robotics can be seen as an important innovation. In contrast to invent yet another path planning algorithm or program a robot control software in C/C++ the open system paradigm reformulates the goals of a robot system. It puts a higher importance on the robot's environment and allows the enviornment to take influence on the robot. There are many examples available in the history of robotics with this background, e.g. Braitenberg vehicle, kismet social robot and SHRDLU. These projects have demonstrated interactive robotics. There is always a robot and a human operator who interacts with the robot.
From a technical perspectives, interactive robotics is equal to teleoperation. Teleoperation was recognized by computer science as opposite to artificial intelligence, because the machine doesn't decide by itself but is guided by external human wisdom. So the maschine can't be called a robot anymore but has more in common with a RC Car.
The rejection of teleoperation makes sense on the first look. If a human operator is in charge to control the RC car, then no artificial intelligence is needed. Therefor it has nothing to do with thinking machines and is located outside of robotics. Only autonomous robots are intelligent robots.
With a modern perspective, Artificial intelligence isn't located inside of a robot but its the interface between a robot and its environment. Such an interface can become smart in the sense that the interface understands natural language.
Practical demonstration of the Total turing test
Stevan Harnad coined around the year 1990 the term "total turing test" which is a philosophical description of an instruction following task in robotics. What is missing is a practical demonstration of such a test for a real robot.
Such a demonstration can be realized in a simple video game modeled as a language game. An entry level example is a navigation task in a graph. There are 8 nodes connected with lines and the robot has to move along the graph to reach a certain goal node. Possible interaction with the robot would be:
- what is your position?
- what is your battery status?
- which nodes are reachable from current position?
- Move north
- move to node #3.
- what is the shortest path to reach node #6?
The robot is in charge to answer these requests in natural language and with motor actions. The problem is easy enough to get implemented as normal computer code without using advanced large language models or vision language action models. The human to robot interaction can be simplified by using a codebook. The amount of possible commands is given in the menu and the human can select one of the commands. IN other words, the Total Turing Test (TTT) is some sort of speaker to hearer interaction game played between a human and a robot.
July 22, 2026
Grounded language in open systems
The box on the left is the human who describes the reality with natural language. The box on the right is the environment which can be perceived with sensors. Symbol grounding is the connection between both boxes.
From a system perspective the 2 box system is an open system because both boxes are connected to each other. Natural language from the left box is referencing to physical objects in the right box, while perceived reality in the right box gets described with English words in the left box.
The assumption is, that there are 2 different systems available which are working with different internal logic. The language layer consists of nouns, verbs, adjectives and grammars which is the symbolic layer. In contrast, the environment has no natural language but it consists of sensor perception, motor actions and 3d objects. The 2 box paradigm describes in a simplified format what natural language is about. Its an abstraction mechanism for the reality. Physical objects like a table or a banana are labeled with words. The ability to label objects is the key element in grounded language and allows to build intelligent robots.
Natural Language as AI technology
Languages likes English or French are discussed by Linguists not by computer scientists. A language is located in the humanities but not within the mathematics department. So its logical that most Linguists have no idea about computer science and vice versa. This might explain why AI research wasn't succesful over decades, because Natural langauge is the missing puzzle piece to make machines intelligent.
From a birds eye perspective, a language like English consists of verbs, nouns and adjectives. The grammar consists of rules how to connect the words to sentences. And language also is directed towards the reality. A word like "a flying bird" or "green flower" is referencing to objects which are available in front of the speaker.
The ability to name every object from the reality with a word and describe activities also with words makes natural language a powerful technology which can be used for human to human communication and machine to human communication both. In case of Artificial Intelligence a computer needs to parse natural language which is the bottleneck in modern AI research. Suppose a computer program asks the user to enter a word. The user enters "red box", then the computer stores the input in a variable but it has no consequences. So the computer isn't able to understand the meaning.
This parsing problem can be solved by inventing a language game. A language game is similar to a 2d arcade game a rule book which explain who to react to a certain input pattern. Typical language oriented games are translation games, the board game Scrabble or the "guess what" game. After implementing these games on a computer, the input of a human will have consequences. The consequences are given by the rules of a certain game, for example in a translation game the user needs to enter the correct translation for a word from another language. The computer verifies if the answer is correct.
Strictly spoken, not the computer decides about the meaning but the rules of the language games are providing the meaning. Modern AI research after the year 2020 is mostly focused on natural language and language games like "Visual question answering", "instruction following" and "question answering chatbots". All these games are formalizing the human to machine interaction in a sense that the computer can determine a score. This score allows to train artificial neural networks. The result of the training process is recognized as Artificial Intelligence in the modern sense. It allows to control robots with language instructions, and generate text with Large language models.
A common assumption in the past was that its very complicated to parse natural language with a computer. This assumption is only correct if the entire corpus of English should be understand by a computer which is around 1 million words and stored in endless amount of books. Parsing this written language with a computer is indeed a hard problem for computer science. What is possible instead is to reduce the task to a subset of English which consists of a dozens words from a restricted domain which are used to play a language game. In the minimal case, there are 6 picture cards and 6 word cards and the task is to match the correct pairs. Such a language game can be implemented in a short computer program and can be played by an automated AI algorithm. The AI algorithm has access to a database with the correct answers. This allows the computer to connect the picture of a banana with the word "banana".
July 21, 2026
Head up display for a kitchen robot
The picture shows an artist version of a head up display. It contains of:
- camera picture of a kitchen
- text box with inner voice
- bounding boxes
- labels for the bounding boxes
Surprisingly, the information in the picture can solve the symbol grounding problem because the head up display connects visual perception with textual information. The text from the inner voice like "I need to find 200g of flour" can be converted into meaning with the help of the bounding boxes. There is a box available with such an ingredient. The task for the robot is not to plan actions but the main problem is to connect language from the inner voice with detected objects in the camera.
Such a link of visual objects with textual labels is the core element in grounded language. If the robot is able to identify objects from the text box, its possible to generate all sort of inner voice. For example, the robot can say that he needs to peel the banana or "open the oven". All these nouns and verbs are translated into position of the bounding box on the screen which allows to execute the action physically.
July 20, 2026
How important is mathematics to understand Artificial intelligence?
On the first look, computer science has its root in mathematics and physics, therefor the assumption is, that AI is based on mathematics too. The precondition, according to the claim, to understand modern robotics is located in analysis, algebra, statistics and boolean algebra.
A closer look will show that mathematics isn't needed in AI research at all or to be more precise, mathematics didn't enable advanced AI in the past. What researchers from 1960 until around 2000s tried was to describe artifciial intelligence in terms of mathematics and logic and they failed. The problem is that none of the mentioned disciplines like statistics, algebra and so on contain a method how to enable artificial intelligence. Even if mathematics is a great science disciplines its complete useless to control a robot. Some attempts were made to solve robotics problems with mathematical description of trajectories including model predictive control, but even these advanced math subjects are not powerful enough to enable autonomous robots.
Artificial intelligence as a science discipline is working different than classic sciences like physics, math and computer science. From a pessimistic perspective there is no such thing like Artificial intelligence. Most research around thinking machines and autonomous robots comes to the conclusion that the promising new technique isn't working. That means, a large scale robot project with hundreds of man year effort isn't working at all, the applied techniques are useless and the researchers have no idea about the cause.
Such a kind of pessimistic situation is not an exception in AI research but it was the default situation from 1960 to 2000s. So the best analogy is to compare AI with a complex puzzle which is impossible to solve and no matter how well the researchers are familiar with mathematics, philosophy or psychology they had no idea how to make a machine intelligent.
AI research in the past was mostly a trial by error meta disciplines which wasn't able to solve any of the goals. It was impossible to build intelligent robots, it was hard to program AI agents for computer games, and speech recognition with a computer was another unsolved topic.
The good news is, that its possible to list techniques who are not leading to AI. These techniques are:
- neural networks
- maschine learning
- reinforcement learning
- mathematical optimiziation
- heuristics
- number crunching
- genetic algorithms
- state space search
- case based reasoning
- expert system
In other words, the entire AIMA book (Russle/Norvig: AI a modern approach) is an anti pattern. It doesn't contain recipes to program robots but it describes the struggle of AI researchers for doing so.
After listing all these techniques who are not leading to AI there is a need to find the shared bias. THis bias is the computational paradigm which means, to use a computer to provide intelligence. This shared bias hasn't worked in the past because its an anti pattern in AI research.
On the first look it makes sense to assume, that AI has to do with computation because the computer or the robot should calculate something which leads to intelligent decision making. Therefor AI has to do with number crunching, programming and mathematics. The problem is that in the reality this paradigm isn't able to control robots but it blocks the progress in technology.
It seems, that AI aka intelligence isn't located inside a robot but outside of the machine. On the first look, such an assumption sounds like blasphemy because outside of a computer there is nothing which can calculate or make decisions. At least until the 2000s such a claim would be rejected by mainstream AI research for sure. With more recent understanding of AI there are some new results available which show that AI might located indeed outside of a robot.
Suppose the source of intelligence is not located in the CPU and not inside of a robot hardware, then intelligence has nothing to do with mathematics or physics. This doesn't mean that intelligence is equal to a magic force but it implies that AI has to do with communication. Communication is the science of how to connect things, communication puts a focus on the air gap between two systems.
The transition from former computational paradigm which locates AI Inside of a robot, towards modern communication paradigm which locates AI between two systems is the major paradigm shift in AI research which took place after the year 2000. The revised understanding of intelligence is strongly connected with communication, linguistics and man to machine interaction. In contrast, former focus on computation, mathematics, algorithms and programming have been discarded.
July 19, 2026
Robot control with head up displays
In contrast to a famous assumption, modern robotics isn't working with algorithms or neural networks but the basic building block is graphical user interface, namely a head up display (HuD). The HuD solves the symbol grounding problem. Typical elements are: bounding boxes around detected objects, text labels for describing the content of a bounding box, another text box for showing the inner voice of rhe robot.
These ingredients are enough to program an advanced artificial intelligence which can solve complex problems. The HuD including the mentioned bounding boxes acts as a communication layer. It ensures that the computer understands basic commands like "move to shelf and grasp the box". A certain high level command is converted into a visual pictures in the HuD, e.g. the word "shelf" is referencing to a bounding box with the label "shelf" which has a 2d position on the screen.
Programming a Head up display for an existing video game is a demanding task but can be solved with standard programming techniques. Most videogames created since the 1980s have a built in debug mode which comes close to a head up display. In the debut mode, all the sprites on the screen are highlighted with frames and sometimes the name of the objects are shown as textual overlay. The combination of graphical display plus textual overlay is the main principle of a head up display and also the main principle of grounded language. So the HuD itself acts as technology for enabling artificial intelligence.
Let me give another example to demonstrate the advantages: Suppose the head up display for a warehouse robot videogame was activated. The user sees some bounding boxes on the screen for highlighting objects in the map like charging station, corridor, shelf A, shelf B, green box, red box. Also the inner voice of the robot is shown a text frame and contains:
"I'm standing at position (3,2). My battery level is 80%, my goal is to fetch the red box from shelf A, the planned trajectory is shown as arrows in the map".
So the initial situation for the robot is, that an annotated HuD is visible which labels objects and mentions the current goal. These information can be translated into actions for the robot. All what the AI of the robot has to do is to compile these information and decide what to do next. From an AI perspective its an instruction following task with an aciivated head up display.
A head up display provides a cognitive space. The shown bouding boxes and labels are creating a symbolic representation of the world. The world of the robot can be described in terms from the head up display. Its no longer a mathematical space and not a 3d space but the reality introduced by the HuD consists of words, locations of items and goals from the inner voice. Such a high level space can be processed by a computer because the amount of possible states is small. There are not millions of possible objects but the HuD shows only 6 different objects in a map. and the inner voice doesn't display millions of possible actions, but the inner voice describes clearly what the current situation is, and what the desired goal state is, similar to a text adventure.




