Robotik und Künstliche Intelligenz ist ein multidisziplinäres Fachgebiet was mehrere hochkomplexe Wissenschaften umfasst wie Naturwissenschaften, Geisteswissensschaften, Sozialwissenschaften und Informatik als verbindenes Element. Dadurch entsteht eine Komplexität und es ist unmöglich in einfachen Worten zu erklären wie ein Roboter funktioniert.
Um all die Wissenschaftsdisziplinen gemeinsam darzustellen im Kontext Robotik, bietet sich ein Head up display an. Ein head up display zeigt die Welt aus sicht eines Roboters und umfasst Kamerafenster und ein textuelles Fenster mit der inner voice des Roboters. Die Softwareentwicklung reduziert sich darauf das Layout und die Funktionsweise des head up displays zu optimieren, also zu bestimmen wo die fenster auf dem bildschirm angezeigt werden, welche Informationen darin zu sehen sind und sicherzustellen dass die daten aktuell sind.
Zum Beispiel werden visuelle Sensoren mit einer hohen Frequenz von 60fps aktualisiert, während die inner voice des Roboters in einer langsamen Frequenz von 1fps aktualisiert wird.
Selbst Laien außerhalb der Künstlichen Intelligenz können anhand eines Head up displays nachvollziehen wie Roboter funktionieren. Sie sehen auf dem Bildschirm was der Roboter sieht und können lesen was der Roboter denkt. Das inner voice textfenster enthält z.B. "Auf dem Tisch steht eine Tasse. Ziel ist die Tasee zu greifen".
Ein Head up display ist das Frontend, im Hintergrund wird weitere Software benötigt wie large language modelle, Vision language action modelle, datasets und bilderkennungssoftware. Diese Softwarebestandteile sind zwar für die funktionsweise nötig haben aber lediglich unterstütztende Funktion. Wichtig ist dass im Head up display alle Softwarekomponenten gemeinsam angezeigt werden.
Das Grundprinzip hinter jedem Head up display ist Bilder und Texte gemeinsam anzuzeigen. Charakterische Elemente sind bounding boxes, textuelle Beschreibungen von erkannten Objekten sowie kurzen Stichworten für Planungsaufgaben. Die Mischung aus Text und Bild ist die Antwort auf das Symbol grounding problem. Es ermöglicht einer Künstlicher Intelligenz Informationen zu speichern. Anstatt das komplette Kamerasignal im speicher abzulegen was sehr viel RAM benötigt, werden ledigliche Textualle Annotationen zu den Kamerabildern abgelegt. Jede Entscheidung, Planung und Ausführung von Aktionen hat etwas mit Verknpfüung von Bild und Text zu tun.
August 19, 2026
Wie man Robotik vereinfacht
Was tun gegen AI slop?
Nach der Dead internet Theory besteht ein großteil des Internet aus KI generierten Inhalten wodurch eine Verflachung einsetzt und menschlich erzeugter Content zurückgedrängt wird. Die logische Frage lautet wie schlimm das ist und wie man den Prozess aufhalten kann.
Leider ist dieser Diskurs über AI slop nicht besonders produktiv sondern führt in eine Sackgasse bei dem technische Details ausgeblendet werden. Das Internet rein auf Inhalte zu reduzieren und die Technologie dahinter als gegegeben zu betrachten ist das eigentliche Problem.
Technisch gesehen entstanden KI erzeugte Inhalte ab dem Jahr 2023 mit dem Aufkommen der ersten LLM welche Text, Fotos, Bilder, Videos und Musik erzeugen können. Die Zeit vor 2023 war das Internet menschlich erzeugt. Computer konnten zwar Inhalte anzeigen in einem Webbrowser und über weite Strecken per Glasfaserkabel übertragen aber Computer waren nicht in der Lage Inhalte zu generieren, von kleineren Ausnahmen wie fraktale und Povray raytracing Bildern abgesehen.
Die Kritik an KI erzeugten Inhalte ist eine versteckte Kritik an large language modellen. Diese sind die technische Basis um mit einem simplen Prompt längere Texte zu erzeugen wie Kurzgeschichten, Fachtexte oder Powerpoint Folien. Historisch gesehen war vermutlich das SCIgen Projekt das erste Beispiel für AI slop. Im Jahr 2005 hat ein Computerprogramm die Inhatle von mehreren Powerpoint Folien erzeugt die dann von verkleideten Wissenschaftlern auf einer fachtagung präsentiert wurden.[1] Aus heutiger Sicht könnte man sagen dass Jeremy Stribling et. al. damals AI Slop erzeugt haben und diesen vor anderen Wissenschaftler präsentiert haben.
Es gibt sogar ein objektives Instrument um AI slop zu erkennen und zwar mit Hilfe eines Turing Tests. Dabei werden einer versuchsperson mehrere Texte präsentiert und die Versuchsperson muss entscheiden ob der Autor ein Mensch oder ein Computer war. Kann die Versuchsperson das nicht anhand der Texte ermitteln hat der Computer den Turing Test bestanden.
[1] https://pdos.csail.mit.edu/archive/scigen/#talks
August 18, 2026
History of the symbol grounding problem
Symbol grounding is a multidisciplinary approach which has evolved over years. The follwoing timeline shows the milestones during the development. First important innovation was the invention of written language thousands of years ago, in the year 1844 the morse code was invented to transmit signs over distance. In 1968, the SHDRLU project was initiated which was a text to graphics interface on a computer. In 1991 Rodney brooks recognized that robots need to interact with the environment, in 2003 a voice controlled robot was built by Deb Roy and in the year 2023 the Lingo-1 Vision language model was published for controlling self driving cars.
The shared similarity between these innovations is the focus on natural language for connecting two systems. A speaker sends a command to a hearer. Implementing such an interface on a computer allows to teleoperate a robot with language. This is the precondition for automation.
The crucial building block for symbol grounding is natural language which is used as a reference system to the external world. Thanks to predefined words for objects and activities its possible to label sensory data and receive high level commands. A robot can recognizhe that he stands in front of an obstacle, and a human operator can submit a commoand to the robot like "move around the obstacle to the left". It took decades and even centuries until human researchers have discovered the power of natural language and were able to implement textual interface for robots. Its likely that future robots invented in 20 years from now will work with the same principle.
3300 BC,Cuneiform writing system in Mesopotamia
1500 BC,sundial showing the time of the day
600 BC,Latin alphabet available in Italy
322 BC,correspondence theory of truth by Aristotle
1386,Salisbury Cathedral tower clock with a bell
1440,printing press by Johannes Gutenberg
1505,Pomander Watch by Peter Henlein
1792,optical telegraph by Claude Chappe
1844,morse code by Samuel Morse
1870,Engine Order Telegraph by William Chadburn
1876,commercial typewriter by Remington
1878,chronophotography "The Horse in Motion" by Eadweard Muybridge
1903,Telekino remote controlled boat by Leonardo Quevedo
1915,Therblig notation by Frank Gilbreth
1915,rotoscoping animation technique by Max Fleischer
1920,AAC Communication board by F. Hall Roe
1928,Labanotation dance notation by Rudolf von Laban
1929,Televox robot by Westinghouse
1930,motion tracking by Nikolai Bernstein
1949,Turing test by Alan Turing
1954,Georgetown-IBM Experiment with russian translation
1959,Pandemonium architecture by Oliver Selfridge
1962,ANIMAC motion capture by Lee Harrison III
1963,ASCII code
1966,ELIZA chatbot by Joseph Weizenbaum
1968,SHRDLU natural language understanding by Terry Winograd
1971,Lexigram for communicating with apes by Ernst von Glasersfeld
1977,Zork I text adventure by Tim Anderson
1977,Tour model instruction following by Benjamin Kuipers MIT AI lab
1980,Chinese room argument by John Searle
1980,Commentator scene description by Bengt Sigurd
1980,Finite State machine in Pacman videogame by Tōru Iwatani
1981,Karel the robot programming language by Richard Pattis
1983,MIDI music protocol
1983,M.I.T. Graphical Marionette by Delle Maxwell
1984,Castle Adventure by Kevin Bales
1987,Maniac Mansion point&click adventure by Ron Gilbert
1987,Vitra visual translator by Wolfgang Wahlster
1989,Speech activated manipulator SAM by Michael Brown
1990,Physical Grounding Hypothesis by Rodney Brooks
1990,paper "The symbol grounding problem" by Stevan Harnad
1991,paper "Intelligence without Representation" by Rodney Brooks
1993,AnimNL computeranimation by Norman Badler
1993,conceptual spaces by Peter Gardenfors
1994,Abigail scene recognition by Jeffrey Siskind
1996,Interaction machines by Peter Wegner
1998,Rocco Robocup commentator by Dirk Voelz
1999,trec-8 Text Retrieval Conference
2000 Kismet social robot by Cynthia Breazeal M.I.T.
2003,M.I.T. Ripley robot by Deb Roy
2006,Marco route instruction following by Matt MacMahon
2007,Simbicon computer animation by Michiel Panne
2010,Motion grammar by Mike Stilman
2010,M.I.T. forklift by Stefanie Tellex
2011,IBM Watson Question answering by David Ferrucci
2013,Word2vec algorithm by Tomas Mikolov
2015,Poeticon++ trajectory recognition by Yiannis Aloimonos
2015,DAQUAR VQA dataset by Mateusz Malinowski
2017,Transformer Architecture by Ashish Vaswani
2020,Vision language model by different authors
2023,Wayve Lingo-1 self driving car
Wie natürliche Sprache Einzug hielt in die Künstliche Intelligenz
Künstliche Intelligenz ist ein sehr traditionsreiches Thema was über Jahrzehnte erforscht wurde. Lange vor der Dartmouth Conference im Jahr 1956 gab es versuche Schacuautomaten und Androiden zu bauen. Auch in der science Fiction literatur waren Roboter ein fester bestandteil, z.B. in der R.U.R. Geschichte von 1920. Obwohl niemand sagen kann wie man denkende Maschinen baut gab es doch verschiedenen Ideen darüber wie sich Roboter technisch realisieren ließen.
Mit der Erfindung des Computers in den 1940er rückten computergesteuerte Roboter in den Fokus. Die wissenschaftlich These lautete fortan dass Künstliche Intelligenz das Resultat von Algorithmen sei, also eines Programs was auf einer CPU ausgeführt wird. Diese Sichtweise prägte über Jahre die Forschung. Bis heute hält sich die Vorstellung unter informatikern, dass Künstliche Intelligenz ein Program sei was von einem Computer ausgeführt wird und Prbbleme lösen kann.
Mitte der 1980er Jahre hat Douglas Lenat das Cyc Projekt ins Leben gerufen. Es war ein umfangreiches Forschungsprojekt was Künstliche Intelligenz nicht über algorithmen realisieren wollte sondern mit Hilfe von Ontologien, also Wissensdatenbanken. Die Idee war dass Intelligenz in semantischen Netzen gespeichert sei. Obwohl das Cyc projekt rückblickend eine Sackgasse darstellte half es dabei eine neue Perspektive zu entwickeln. Anders als sortier oder optimierungsalgorithmen benötigen Ontologien keine CPU sondern werden auf Festplatten gespeichert, sind also Datenstrukturen.
Es dauerte noch einige Jahre bis das Ontologie konzept erweitert wurde zur These dass Künstliche Intelligenz etwas mit natürlicher Sprache zu tun hat. Anders als die Ontologien aus dem Cyc Projekt lässt sich natürliche Sprache sehr viel schwerer von Computern parsen. Sprache ist nicht als graph organisiert sondern es gibt Wörter die eine BEdeutung haben die aber nicht im Satz selber enthalten ist. Um natürliche Sprache dennoch mit Computern zu verarbeiten benötigt man weitere Hilfsmittel wie datasets in welchen Interaktionen mit Sprache gespeichert sind. Diese Datasets dienen als input für neuronale Netze womit sich LLM Chatbots erzeugen lassen die aktuell im Jahr 2026 die KI Landschaft prägen.
Rein technisch hätte bereits in den 1960er Jahren ein Informatiker behaupten können dass künstliche Intelligenz etwas mit natürlicher Sprache zu tun habe. Und es wäre möglich gewesen auf damaligen Computern ein Bilderkennungsprogram schreiben zu können was ein Bild in Sprache konvertiert. Das problem war, dass die Idee viel zu fortschrittlich war als dass sie in den 1960er geläufig wäre. Selbst heute gilt die Vorstellung dass Computer natürliche Sprache verwenden so wie Menschen als provokative Äußerung. Sprache wird zwar von Philosophen untersucht und von Linguistikern in Wörterbüchern explizit gespeichert, aber der Konsens ist dass nur Menschen sprechen nkönnen, nicht aber Tiere und erst recht nicht Computer.
Die Verwendung von Sprache als Technologie um künstliche Intelligenz zu ermöglichen war ein langer prozess innerhalb der Informatik der Jahrzehnte benötigte. Zu nennen ist beispielsweise der M.I.T. Ripley robot von Deb Roy in 2003, oder der Word2vec Algorithmus aus 2013. Technisch gesehen sind beide Projekte höchst unterschiedlich aufgestellt, einmal geht es um die Steuerung von Robotern und beim zweiten Beispiel um präprocessing für neuronale Netze. Die Gemeinsamkeit besteht darin, dass natürliche Sprache als zentrales Element für ein KI Projekt eingesetzt wird. Demnach ist jenes Modul womit der Computer denkt, zugleich das Modul was sprache versteht und Sprache generiert.
Sprache ist deswegen der zentrale Baustein von künstlicher Intelligenz weil damit eine Referenz auf die Umwelt des Roboters erfolgt. Die Einbettung des Roboters in eine Umwelt wurde zuerst von Rodney Brooks in den 1990er Jahren erkannt. Es ist nicht ausreichend eine Maschine zu bauen und darauf ein Programm auszuführen sondern Robobter müssen über sensoren die Umwelt wahrnehmen. Die Intelligenz wird also erst über eine Interaktion mit der Umgebung erzeugt, und Sprache ist dafür das Interface. Sobald ein Roboter Sprache parsen und erzeugen kann besitzt der Roboter über ein leistungsfähiges referenzsystem. Er kann auf Kommandos reagieren, er kann sensor signale in Sprache umwandeln, er kann beobachtungen in einer Log datei speichern. Daduruch entsteht Künstliche Intelligenz.
August 17, 2026
Wodurch werden KI Projekte so teuer?
Der Diskurs über Deutschland als Technologiestandort wird häufig verengt zu der frage warum es keine KI Forschung in Deutschland gibt. Dabei gab es zumindest in der Vergangenheit viele deutsche KI Projekte wie:
- verbmobil, Vitra, RHINO (Museumsroboter), ARMAR (humanoider Roboter), RoboCup@Home, Carolocup und Dickmanns autonomes fahren
Das Grundproblem bei KI Forschung ist dass es extrem teuer ist im Sinne von finanzieller Kosten. Ende der 1980er benötigten Forscher die schnellste verfügbare Hardware um darauf Lisp laufen zu lassen und erste neuronale Netze zu simulieren wie VAX Minicomputer (0.5 Mio US$ pro Stück) und Symbolics Workstation (100k US$ pro Stück). In diesen Computern war für die damalige Zeit sehr viel RAM verbaut und die Systeme waren Einzelanfertigungen aber kein MAssenprodukt wie der damals verfügbare Atari ST Rechner.
Und das sind nur die Ausgaben für die Hardware, hinzu kommen die Kosten für das KI Projekt selber. Künstlicher Intelligenz unterscheidet sich von klassischer Forschung dadurch dass es interdisziplinär betrieben wird. Neben der Informatik braucht man zugriff auf Mathematik, Sprachverarbeitung, Kognitionspsychologie usw. Allein eine Bibliothek welche all diese Literatur bereitstellt ist ein hoher Kostenfaktor.
Sowetwas wie preiswerte KI forschung welche auf Standardhardware läuft und nur ein Wissenschaftsgebiet umfasst gibt es nicht. Das wäre dann z.B. Forschung innerhalb der Mathematik oder Forschung innerhalb der Mechanik. Bei künstliche Intelligenz fließen all diese Bereiche zusammen. Ein robotik projekt besteht aus hardware, software, Algorithmen, Natürliche Sprache, Bilderkennung, neuronalen Netzen, Dateübertragung über Netze usw.
Bei KI Forschung im Jahr 2026 haben sich die Kosten weiter erhöht. Aktuelle Hardware um neuronale Netze zu trainieren ist sehr teuer, die Notwendigkeit unterschiedliche Disziplinen einzubeziehen hat sich verstärkt. Roboter werden heute über motion capture datensätze trainiert welche zuvor in Laboren für Bewegungsstudien erfasst wurden, DIe sprachverarbeitung erfolgt über Datenbanken in denen fremdsprachichige Korpora gespeichert sind und Bilderkennung benötigt riesige Beispieldatenbanken mit hunderten von Terabyte an Speicher. All diese Ressourcen sind teuer, die akademische Publikationen liegen geschützt hinter kostenpflichtigen Paywalls, selbstfahrende Elektroautos mit Lidar sensoren kosten deutlich mehr als ein Standard Auto und es dauert Jahrzehnte bis sich nachwuchsforscher in die Thematik eingearbeitet haben.
Grundlagen forschung im Jahr 1900 in den Bereichen Funkübertragung oder Elektronik war relativ einfach durchzuführen. Es braucht nicht mehr als einen Bastelkeller, einige wenige Grundlagenwerke aus dem Fachgebiet, ein wenig praktisches geschick beim Aufbau von Schaltungen und schon konnte die Spitzenforschcung beginnen.
Die schlechte Nachricht lautet dass es nicht möglich ist die Kosten zu senken. Neuronale Netze auf älterer Hardware zu trainieren funktioniert technisch nicht, Robotik ohne Sprachverarbeitung zu erforschen macht inhaltlich keinen Sinn, Trainingsdateaets zu erzeugen die nur wenige Megabyte groß sind erzeugt schlechte Resultte, und humanoide Roboter zu bauen die nur 3 servomotoren haben und über ein Stromkabel versorgt werden ist kein guter Standard in der autonomen Robotik.
Es gibt jedoch eine Methode um die Kosten massiv zu senken. Und zwar indem man KI Projekte rein virtuell durchführt. Das bedeutet, es gibt keine physische Hardware und es gibt keine realen Sensoren sondern die KI agiert in Computerspielen. Einsteigerfreundliche Projekte wären:
- Steuerung eines Pong Clones mit Hilfe neuronaler Netze
- Programmierugn einer SChach Engine
- simulation eines 6 beinigen Laufroboters als 3d Modell
Der verbleibende Kostenfaktor bei diesen Projekten ist die inhärente Multidisziplinarität, das also unterschiedliche Wissenschaftliche Bereiche wie Mathematik, Informatik, Biologie, Sprachwissenschaft, Statistik, Bewegungsstudien miteinander kombiniert werden. Möchte man z.B. einen simplen 6 beinigen Laufroboter in 3d Animieren benötigt man für diese Aufgabe unhzählige Bücher aus mehreren Bereichen der Unibibliothek. Es gibt kein Fachgebiet in einer Bibliothek was sich dezidiert mit laufrobotern beschäftigt, sondern es gibt nur Fächer die etwas über Mechanik, Programmierung, Animation, Algorithmen und Bewegung beinhalten.
Das was unter dem Stichwort Künstliche Intelligenz gemeint ist, stellt ein Mix da aus all diesen Disziplinen. All diese Informationen zu lesen und zu kombinieren ist zeitaufwendig. Es dauert lange bis sich newbies in die Thematik eingelesen haben.
Mit Hinblick auf den Erfolg von Large language modellen ab dem Jahr 2020 ist zwar klar wie erfolgreiche KI projekte aussehen, allerdings sind Large language projekte gleichzeitig die aufwendigsten KI Projekte überhaupt. Man braucht zwingend große datasets im Terabyte Umfang und schnelle GPU für das Training der neuronalen Netze. Insofern ist die Hürde solche Projekte zu beginnen sehr hoch.
In den 1980er Jahren kosteten Großforschungsprojekte im Bereich Künstliche Intelligenz 50 Mio DM. Aktuelle Forschung im Jahr 2026 mit Large language modellen kostet rund 500 Mio US$. Vermutlich werden künftige KI Projekte bei denen Robotik mit large language modellen kombiniert werden, Kosten von 5 Millarden US$ erzeugen. und sobald man damit beginnt Schwärme von Robotern zu bauen wird es nochmals teurer.
August 11, 2026
Einführung in language games
Language games, deutsch: Sprachspiele nach Wittgenstein, sind eine besondere Form des Gesellschaftsspiel. Anders als das bekanntere Schachspiel wurde es in der Geschichte der Künstliche Intelligenz lange Zeit nicht näher untersucht. Was hingegen von der KI Community sehr intensiv erforscht wurde waren Spiele wie Schach, Mühle, das Piano movers problem im Kontext Motion Planning sowie Roboternavigation in einem Labyrinth. Diese klassischen Spiele werden in der Literatur diskutiert und es gibt unzählige Software mit denen einen Künstliche Intelligenz die Spiele gewinnen kann.
Die eingangs erwähnten Sprachspiele nach Wittgenstein sind selbst innerhalb der Philosophie ein Randgebiet. Es gibt dazu zwar Publikationen aber nur wenige und diese sind gänzlich theoretischer Natur. Im Kontext des Symbol grounding problems werden Language games jedoch zu einem wichtigen Werkzeug zur Erklärung von Mansch maschine interaktion. Ein Sprachspiel ist zunächst einmal ein Gesellschaftsspiel was nach Regeln abläuft. Dazu ein Beispiel:
Das wohl einfachste verfügbare Sprachspiele ist "Farben raten". Der Spielleiter zeigt eine farbige Karte und der Spiel muss das passende Wort sagen, z.B. "blau". dann zeigt der Spielleiter ein andere Karte und der Spieler muss erneut das passende Wort sagen, z.B. "hellgrün". Am Ende wird die Punktzahl ermittelt, also wieoft der Spieler das richtige Wort gesagt hat.
Wenn man das Spiel in seiner Muttersprache spielt ist es trivial, deutlich schwerer wird es hingegen in einer Fremdsprache. Der Spieler sieht zwar dass die Karte "hellgrün" ist kennt aber das passende Wort in der Fremdsprache z.B. italienisch "verde chiaro" nicht.
Es gibt neben "Farben raten" noch weitere Sprachspiele wie "instruction following", "Name guessing" die ebenfalls nach festen Regeln gespielt werden und etwas mit dem Aussprechen von Worten zu tun haben und die Wirklichkeit zu benennen. Sprachspiele werden praktisch im Fremdsprachen unterricht eingesetzt um neues Vokabular in einer realistischen Situation anzuwenden. Es ist eher unüblich, Sprachspiele in der Informatik in Software zu implementieren, jedenfalls ist die Menge an Litertur zu dieser Thematik fast null.
Es spricht technisch nichts dagegen das obigen Farben raten und weitere language games in Software zu implementieren sowie KI Agenten zu programmieren die diese Sprachspiele lösen können. Dies wird meist als Vision language action model bezeichnet, also eine KI die die Realität in Sprache beschreibt und darauf reagieren kann.
Language games sind in der Informatik ein sehr mächtiges Werkzeug, selbst wenn man sehr simple Sprachspiele implementiert die aus wenigen Worten bestehen kann der so instruierte Roboter hochkomplexe Aufgaben lösen. Scheinbar sind Sprachspiele die Kernkomponente von künstlicher Intelligenz. Vielleicht ein Beispiel: angenommen man überlegt sich ein Sprachspiel für Küchenroboter. Nachdem das Sprachspiel in softwrae implementiert wurde und der Roboter Begriffe der Küche korrekt bennent, kann dieser Roboter bereits längere Aufgaben ausführen. Er versteht plötzlich eine Anweisung wie "öffne den SChrank und entnehme die Tasse". Dieses Kommando wird deshalb verstanden weil für den Roboter es Teil eines Sprachspiels ist. Die Aufgabe ist die Worte in dem Satz in der Realität zu suchen, z.b. über bounding boxes.
Sprachspiele haben einen Schwerpunkt in der Kommunikation. Anders als beim piano movers problem geht es nicht um Motion planning sondern bei Sprachspielen geht es um assoziatives Wortverständnis, also das matchen von Begriffen mit Bildern. z.b. zeigt der Spielleiter ein Photo von einer Tasse und der Spieler muss das korrekte Wort in Deutsch sagen "Tasse". Es werden also Fähigkeiten abgefragt die etwas mit Linguistik zu tun haben aber nicht mit Mathematik oder Zahlenverständnis. Dies könnte erklären warum Sprachspiele von der Informatik nur selten thematisiert werden, weil die Informatik sich historisch als Erweituerung der Mathematik versteht, es geht darin weniger um Texte oder Worte sondern der Untersuchungsgegenstand sind Zahlen und Algorithmen. Das könnte erklären warum früher in der KI Forschung überwiegend Spiele untersucht wurden die einen mathamtischen hintergrund besitzen also z.B. TicTacToe, Sudoko Spiele, oder das Nim spiel, was eines der ersten Spiele überhaupt war, das auf einem Computer implementiert wurde und zwar 1940 während der Weltausstellung in New York.
DIKW pyramid as blueprint for Artificial Intelligence
A DIKW pyramid divides facts into separate layers:
- low level = sensory data, e.g. color=(20,10,178), distance=2 meter
- mid level= label space, e.g. [bluecolor], [mediumdistance], [baterryfull]
- high level = knowledge. e.g. "bring me the blue box from room A"
Its not an algorithm but a data representation format, similar to the Unicode table. After implementing the DIWK pyramid as a computer program its possible to interact with the robot in natural language. The human operator formulates a request on the high level layer, e.g. "bring me the blue box" and this request is translated into mid level and low level.
The principle is similar to the TCP/IP protocol layer which allows divides communication into layers for reducing complexity. The working thesis is, that a dikw pyramid can explain what artificial intelligence is about. If the robot has access to a DIKW pyramid, the robot is intelligent.
The main purpose of a DIKW pyramid is to act as an interface between internal world and external world. The low level layer is located inside the robot, while the high level layer is located outside of the robot. Communication means to transmit messages physical but also translate the message between the layers.
August 09, 2026
Grenzen autonomer Robotik
Der simple Grund warum bis ca. 2010 es unmöglich war Roboter zu bauen und zu programmieren liegt darin, dass unklar blieb wie genau komplexe Aufgaben in Software algorithmisch bewschrieben werden sollten. Wenn z.b. ein Roboter den kürzesten Weg zum Ziel finden soll um dort einen Gegenstand präzise zu greifen benötigt der Roboter eine extrem komplexe Steuerungssoftware aus vielen tausend Zeilen von Code. Dieser Code muss irgendwer programmieren.
Soll der Roboter komplexere Aufgaben lösen z.b. bei selbstfahrenden Autos oder beim biped walking erhöht sich der Programmieraufwand weiter. Zwar konnten KI Forscher bis 2010 grundsätzlich den benötigten C/C++ Softwarestack programmieren udn mögliche Fehler darin korrigieren, nur eben in einem sehr langsamen Tempo mit den üblichen 10 lines of code pro Tag für neu zu erstellende Software. Es verwundert wenig das größere Robotik-Projekte mit diesem Entwicklungstempo unmöglich zu realisieren waren.
Die Antwort auf das Dilemma besteht keineswegs darin über genetische Algorithmen die Software sich selbst verbessern zu lassen sondern die benötigte Technologie um Roboter zu programmieren besteht in der Externalisierung von Intelligenz. Je mehr Aufgaben der Roboter an eine Umwelt deligiert desto schlanker ist die verbleibende Robotik-Software. Ein anderer Begriff für Externalisierung lautet Teleoperation. Das heißt die Wegfindung oder die Greifplanung eines Roboters wird nicht länger in Software realisiert sondern durch einen menschjlichen Bediener außerhalb des Roboters.
Anstatt das Für und Wieder von Teleoperation zu untersuchen ist es wichtiger sich die Sprachinterfaces zu kümmern die zwischen dem Roboter und dem menschlichen Bediener benötigt werden. Dieses Interface funktioniert in bidirektionaler Richtung: der Roboter reagiert auf gesprochene Kommandos und der Roboter sendet verbale Statusinformation an den Menschen zurück.
Mit Hilfe eines natürlichsprachlichen Interfaces lassen sich einerseits hochkomplexe Roboter realisieren wie z.b. selbstfahrende Autos, während gleichzeitig die benötigten Codezeilen gering sind. Die Aufgabe der Robotersoftware ist weniger das Steuern des Roboters, sondern das Ziel ist ein Sprachinterface bereitzustellen damit der Roboter effektiv aus der Ferne gesteuert werden kann.
Bis ungefähr zum Jahr 2010 erschien das Konzept von ferngesteuerten Robotern absurd. Es wurde nichtmal kritisch distanziert von der KI Community diskutiert sondern überhaupt nicht thematisiert. Der Konsens stattdessen lautete3 autonome Roboter zu bauen die ohne Teleoperation funktionieren. 100% aller Robotikwettbewerbe wie Robocup, micromouse, Carolo-Cup oder die Darpa Grand challenge arbeiteten nach diesem Prinzip.
August 08, 2026
Sprachspiele im Kontext von Robotik
Die Erforschung künstlicher Intelligenz drehte sich lange Zeit um die Frage welche Art von Softwareframework, Algorithmus oder Programmiersprache benötigt wird um denkende Maschinen zu realisieren. Die Annahme hinter dem Cyc Projekt von Douglas Lenat lautete, dass eine Ontologie der Grundbaustein sei, das Cam-Brain Projekt von Hugo de Garis unterstellte dass im Kern ein neuronales Netz benötigt wird während Edward Feigenbaum vermutete dass sich Künstliche Intelligenz mit Hilfe von Expertensystemen realisieren läst.
Im Laufe der Zeit wurden sehr viele gegensätzliche Technologien und Annahmen entwickelt um Künstliche Intelligenz zu verwirklichen. Ein Sprachspiel im Sinne von Ludwig Wittgenstein ist nur ein weiterer Vorschlag unter vielen. Dennoch lohnt es sich, das Thema nähter zu untersuchen. Weil das Prinzip von Sprachspielen deutlich anders ist als z.b. ein Expertensystem oder ein neuronales Netz.
Sprachspiele sind ein formalisiertes Interface zwischen interner und externer Realität von Systemen. Es ist ein Test, ob die Spielteilnehmer Sprache erzeugen und verstehen mit derer man die externe Realität beschreibt. Ein typisches Beispiel ist das "Ich sehe was du nicht siehst Spiel". Dabei beschreibt Person A ein Objekt aus der Realität anhand von Eigenschaften, er formuliert aussagen wie "das Objekt ist rund, steht in der Küche, hat Beine" und daraus folgert Person B "es ist ein Tisch".
Die Gemeinsamkeit von allen Sprachspielen ist, dass Sprache in einem interkationsspiel genutzt wird um auf die Wirklichkeit zu referenzieren. Es geht immer darum, Dinge abzufragen die in der Realität vorkommen, Anweisungen geben was in der Realität zu tun ist oder sonstwie Sprache und Wirklichkeit in Beziehung zu setzen. Sprachspiele sind eine gute Möglichkeit eine Fremdsprache zu erlernen und dienen dazu den Wissensstand zu erfassen. Wenn eine Person oder ein Computer in einem Sprachspiel eine hohe Punktzahl erreicht, ist diese Person mit der jeweiligen Sprache vertraut.
Anders als die eingangs erwähnten Technologie zur Realisierung von künstlicher Intelligenz wie Onotologien oder Expertensysteme sind Sprachspiele nicht technisch definiert sondern haben ihren Ursprung in der Philosophie und der Linguistik. Es geht um Themen wie interaktivität, Sprache, Wirklichkeit. Die Annahme lautet dass innerhalb dieser Begriffe Künstliche Intelligenz möglich ist, das also denkende Maschinen ein Interface zwischen internem System und externer Realität sind.
Seit den 1980er wurden Brettspiele als Testumgebung für denkende Maschinen genutzt. Schach ist ein sehr altes Beispiel für das Künstliche Intelligenz realisiert wurde, aber auch andere Spiele wie Tic Tac Toe, Dame, Backgammon und Go sind klassische Umgebung zur Erforschung von KI Algorithmen. Leider haben diese Brettspiele den Nachteil dass sie nicht gut nach oben skalieren. Ein Computer der perfekt Schach spielt ist nicht automatisch im Stande einen Roboter zu steuern. Deshalb eignen sich die erwähnten Brettspiele nur sehr eingeschränkt dazu KI näher zu erforschen.
Sprachspiele kann man als neuartiges Gesellschaftsspiel verstehen was ähnlich wie Schach Regeln folgt aber viel besser nach oben skaliert. Ein Computer der das "Guess what" Sprachspiel beherscht ist zugleich auch in der Lage mit diesem Wissen einen Roboter zu steuern. Scheinbar können Sprachspiele die Kernidee von Künstlicher Intelligenz viel besser formalisieren als frühere Brettspiele.
August 07, 2026
Old school artificial intelligence
Before the advent of large language models, there was lots of AI related research available which was mostly ignored by the public. Typical topics in robotics from the past was NEAT neuroevolution, model predictive control and genetic algorithms. All these concepts were working with computational paradigm. At first the fitness function was defined, e.g. the robot gets +1 reward for moving on step forward, and then an algorithm modifies the parameters of the neural network. In case of model predictive control it was even possible to plan one step ahead, so that the robot was able to use a trial&error strategy.
These concepts can be called old school artificial intelligence because they were working with computers and algorithms as the core element. The assumption was that artificial intelligence is equal to search, so its a mathematical optimization problem. Instead of finding the prime numbers from 1 to 1 million and find 1000 digits of the PI constant, the goal was to use raw processing power to solve robotics problems.
The mentioned NEAT and genetic algorithms strategy were the core principle of AI in the past. Large scale projects like CAM brain and the fifth generation computers in Japan were built around these subjects. The researchers have treated AI problems as mathematical optimization problems by transforming robot motion planning into a mathematical state space first, and then use algorithms to solve these problems.
What was unclear until around 2010 was, that all these strategies have failed. Mathematics and algorithms in general are not powerful to solve robotics problems. This insight is even today some sort of blasphemy because mathematics and algorithms are the building blocks in computer science. If both concepts are rejected as weak then the amount of remaining strategies is small and even empty.
The major result of old school artificial intelligence was not to build robots but to proof that mathematical optimization is a dead end. It can be formulated from a theoretical perspective, that all interesting robotics problems are np hard, and that genetic algorithms including NEAT are not able to solve this problem category. Therefor AI can't be realized.
Its possible to replicate the failure of NEAT and similar approaches with today's software libraries. At first a robotics problem gets formulated, e.g. a walking robot in the OpenAI gym which is a 2d simulator. then a state of the art neuroevolution library is utilized to control the robot. A handcrafted fitness function will reward the algorithm so that the optimization can evolve into improved biped walking generator. Unfortunately, the described setup won't work in reality. Even if the robot walks some steps forward the robot will struggle on larger obstacles. At the same time, the CPU consumption is very high.
At first the fitness function of the robot will improve, but then there is no further improvement available and its impossible to fix the problem. Some reinforcement learning advocates are explaining that with parameter tuning the problem can be solved, other NEAT users claim that with future algorithm update the issue might be solved. But the sad reality is, that mathematical optimization in general fails for robotics problem. No matter if NEAT, genetic algorithm or Q learning was utilized, none of these strategies can control a biped walking robot.
Such kind of insight isn't the end of robotics research but it provides meaningful facts about the limitation of old school algorithms. It shows on a practical example, what sort of AI algorithm isn't working in reality. This opens the path for more advanced AI techniques beyond classical computer science.
August 03, 2026
KI als Berechnungsprozess
Zumindest früher und in einem universitären Kontext wurde Künstliche Intelligenz erstaunlich präzise definiert. Rückblickend betrachtet war diese Definition eine Sackgasse welche paradoxerweise das Entstehen von leistungsfähiger Robotik sowie von Large language modelle verhindert hat anstatt zu ermöglichen. Dennoch ist es wichtig sich mit dieser philosophischen Sackgasse näher auseinanderzusetzen.
In einem älteren Foliensatz aus dem Jahr 2005 wurde erläutert dass Künstliche Intelligenz ein Berechnungsprozess sei. [1] page 3. Um ein Problem zu lösen müsse man im Problemraum suchen. [1] page 5
Es werden mehrere Beispiele genannt wie das "Affe und Banana Problem" und "Der Turm von Hanoi" und es werden Lösungsalgorithmen vorgestellt wie uniformierte Suche, Backtracking und Hill climbing mit Bewertungsfunktion.[1] page 12
Insgesamt ist die Darstellung richtig und für die damalige Zeit state of the art. Foliensätze von anderen Universitäten waren ähnlich aufgebaut. Die Grundannahme lautet dass ein klar definiertes Problem existiert was sich über Suchalgorithmen lösen lässt. Als bestes Suchverfahren gilt Hill climbing mit Bewertungsfunktion was auch als heuristische Suche bezeichnet wird.
Die Aufgabe des Computerprograms besteht darin diese Suche auszuführen. Also eine Berechnung durchzuführen welche als Algorithmus definiert wurde. Demnach ist KI die Anwendung einer CPU um Probleme zu lösen.
Was zumindest in der Vergangenheit nicht bekannt war oder verdrängt wurde ist, dass das Verfahren nur bei sehr simplen Problemen funktionoiert wie Finden eines Weges in einem Labyrinth oder Computerschach. Sobald die Probleme geringfügig anspruchsvoller werden, z.B. sehr große Labyrinthe, komplexere Brettspiele oder gar Robotik versagt das Verfahren. Ein typisches Beispiel ist Greifplanung für eine Roboterhand. Selbst wenn eine Bewertungsfunktioni existiert gelingt es nicht mittels Suche im Problemraum eine Lösung zu ermitteln. Der Suchraum ist viel zu groß.
[1] Joachim Funke: Seminar: Denken und Problemlösen, Ruprecht-Karls-Universität Heidelberg, 2005, https://www.psychologie.uni-heidelberg.de/ae/allg/lehre/050217_Computermodelle.pdf
August 01, 2026
KI entwickeln ohne die Verwendung Künstlicher Intelligenz
Während der 1980er Jahre versuchte die KI Forschung Computern das Denken beizubringen. Vorhandene Programmiersprachen wie C oder Pascal wurden als nicht leistungsfähig genug betrachtet und durch dezidierte 5th generation languages wie Prolog und Expertensystems shells ersetzt. Die Zielstellung war, vorhandenes Wissen leichter auf einen Computer zu übetragen damit anschließend der Computer das Problem zu lösen vermag.
Einen ähnlichen Ansatz verfolgt das Semantic Web. Auch dort wurde die komplexe Wirklichkeit in Fakten gespeichert und dann von Rule-Engines ausgewertet. Man kann also semantic web als die logische Weiterentwicklung einer Expertensystem shell betrachten.
Im wesentlichen ging es darum, Teile der Wirklichkeit in ein Computerprogramm zu konvertieren. In diesem Computerprogramm konnten dann logische Schlüsse gezogen werden, eine Suche durchgeführt werden und Entscheidungen getroffen werden. Die Sprache Prolog ist sehr gut dafür geeignet, ein logisches Wissensnetzwerk zu erzeugen. Man gibt zuerst die Fakten in das Program ein und kann dann fehlende Informationen daraus folgern.
Die Stillschweigende Annahme hinter diesen Technologien bestand darin, dass eine KÜnstliche Intelligenz zwingend auf einem Computer ausgeführt wird in Form einer Software. Beim Semantic web gab es eine Wissensbasis gespeichert im Computer plus eine Rule engine, bei der Prolog Sprache gab es die Software einerseits und den Prolog interpreter andererseits. Diese strukuturelle Kopplung aus Künstlicher Intelligenz und Informatik ist historsich gewachsen, nur leider war sie zugleich für den KI Winter Anfang der 1990er Jahre verantwortlich.
Solange man die Zielstellung verfolgt, ein Computer programm zu erstellen oder eine Wissnsbasis anzulegen die von einem Computer program ausgewertet wird entsteht ein möglicher Reality gap. Also der Unterschied zwischen der Wirklichkeit und dem Computermodell was diese Wirklichkeit simuliert. Die einzige Methode diesen gap zu minimieren besteht darin, die Anzahl der regeln im Prolog Program zu erhöhen. Dadurch steigt die Komplexität der Software.
Die Antwort zum Reality gap problem lautet: Fernsteuerung. Es gibt also keine Robotik-Software welche den Roboter steuert sondern es gibt einen menshclichen Bediener. Und es gibt auch kein Expertsystem was eine Diagnose erstellt sondern ein menschlicher Experte erstellt die Diagnose. Durch die simple Forderung nach Fernsteuerung löst man alle Probleme der Künstliche Intelligenz der 1980er. Man tauscht das vorherige Berechenbarkeitsparadigma aus zugunsten eines Kommunikationsparadigmas. Der Computer ist nicht mehr dafür zuständig Programcode auszuführen, sondern der Computer übertragt Funksignale von A nach B.
Diese veränderte Rollenzuweisung entspricht weitaus besser der technischen Leistungsfähigkeit. Wenn die einzige Aufgabe des Computers darin besteht Signale weiterzuleiten sind die benötigte Hardware- Software- und Algorithmen trivial. Gleichzeitig entsteht ein neues Dilemma: Fernsteuerung bedeutet, dass kein Computer die entscheidung trifft sondern ein Mensch. Das wird üblicherweise nicht als KI betrachtet. Sobald ein Mensch mit Fernsteuerung den Roboter steuert, hat man eben keinen autonomen Roboter mehr der über eine Software Entscheidungen trifft sondern man hat ein normales RC Car realisiert.
Erst mit einem erweitereten KI Begriff der distributed cognition mit einschließt wird ein ferngesteuertes Auto zu einer Künstlichen Intelligenz umdefiniert. Distributed cognition ist eine relativ neue philosophische Richtung die davon ausgeht, dass Intelligenz sich aus der Kommunikation der Invididuum ergibt. Demnzufolge kann eine einzelne CPU in einem isolierten Roboter nicht intelligent sein, erst wenn diese CPU mit einer anderen CPU oder einem Menschen interagiert, entsteht Intelligenz.
July 31, 2026
KI Definition in der Vergangenheit
Eine häufig formulierte Definition von Künstlicher Intelligenz früher lautet, dass es darum geht menschliches Denken mit einem Computer zu simulieren. Diese Definition ist abgeleitet vom Turing Test, kann aber auch ohne einen mathematischen Hintergrund formuliert werden. Auch außerhalb der Informatik ist ein Roboter eine Maschine die so aussieht wie Mensch und das selbe leistet. Es geht also darum das menschliche Denken in einem Computer nachzubauen.
Als Einstieg mag diese Definition nützlich sein, allerdings gibt es das Problem dass unklar bleibt wie genau ein Computer einen Menschen nachbilden soll. Computer werden über Software programmiert und diese Software muss sehr präzise Angaben enthalten was der Computer tun soll. Genau an dieser präzisen Vorstellung bezüglich menschlichen Denkens fehlt es. Es gibt zwar durchaus Ansätze den Menschen zu untersuchen mit Hilfe von Biologie, Psychologie und Soziologie nur lassen sich die dabei gewonnenen Erkenntnisse nicht in Software übertragen. Als Folge wird Künstliche Intelligenz zu einer theoretisch philosophischen Betrachtung verkürzt und kann nicht durch Experimente falsifiziert werden.
Die Definition künstlicher Intelligenz bildet den Rahmen innerhalb derer die Erforschung vorangetrieben wird. Die obige Definition "KI=Menschen nachbauen" ist zu ungenau. Dadurch wird KI verunmöglicht.
Hier ist eine modernere Definition von KI, welche den Fokus auf Interaktion mit der Umwelt richtet:
"Künstliche Intelligenz ist die Fähigkeit eines Systems, in Echtzeit mit seiner Umgebung (einschließlich Menschen) zu interagieren, indem es sensorische Eingaben verarbeitet, Handlungen plant und ausführt sowie durch Feedback (z. B. Sprache, Gesten oder andere Signale) sein Verhalten anpasst, um zielgerichtete Aufgaben zu erfüllen oder Kommunikation zu ermöglichen."
July 29, 2026
KI als Mapping problem in der DIKW pyramide
Zumindest bis in die 2000er Jahre wurde Künstliche Intelligenz so definiert, dass menschliches Denken mit Hilfe eines Computers simuliert werden soll. Aufbauend auf dieser Definition wurden mehrere Strategien und Algorithmen diskutiert, die jedoch das selbst definierte Ziel nicht zu erreichen vermochten. Der wohl erfolgreichste Algorithmus der KI Forschung bis zum Jahr 2000 war der bekannte Minimax Algorithmus um Schach von einem Computer spielen zu können. Dieser Algorithmus war so erfolgreich, dass er sogar Großmeister schlagen konnte. Nur, Minimax lässt sich nicht auf Robotik-Probleme verallgemeinern.
Eine neuere und weniger bekannte Definition von Künstlicher Intelligenz lautet dass es ein Zuordnungsproblem ist zwischen low level daten und high leven natürlicher Sprache. Dieses Zuordnungsproblem entsteht innerhalb der DIKW pyramide (data, information, knowledge, wisdom). Vielleicht ein kleines Beispiel:
Über Motion capture wird die Bewegung eines menschlichen Aktors aufgezeichnet. Die x/y/z position der Mocap Marker sind Daten und gehören zum Data layer der DIKW pyramide. Jetzt soll über Mustererkennungsverfahren ermittelt werden, welche Pose der menschliche Aktor gerade ausführt. Diesre kann gehen, rennen, sitzen oder springen. Diese Labels werden im Information layer der DIKW pyramide gespeichert. Aufgabe für die Künstliche Intelligenz ist es jetzt zwischen beiden layern eine Zuordnung herzustellen. Es geht also um die Frage wie eine mathematische Realität auf eine linguistische Realität projiziert werden kann.
Die naheliegende Frage lautet: Warum muss der mathematische Raum der Mocap Marker auf einen linguistischen Raum der textuellen Label projiziert werden? Kann man die Steuerung von Robotern nicht auch eleganter/einfacher beschreiben?
Die Notwendigkeit der Sprachlichen Enkodierung ergibt sich aus dem Kommunikationserfordernis. Es reicht nicht, die mocap Marker als x/y/z position zu speichern, sondern zusätzlich soll einem außenstehender Beobachter in natürlicher Sprache übermittelt werden, was genau der Aktor gerade tut. Ziel ist weniger ein System intern zu beschreiben, sondern Zielstellung ist mit einem 2. System zu kommunizieren.
KI Forschung bis ca. zum Jahr 2000 war geprägt von der closed system Hypothese. Ziel war stets das interne Funktionieren eines Computers zu optimieren durch bessere Algorithmen, schnellere Programmiersprachen oder spezielle Datenbanken wie Cyc. Im Gegensatz dazu verfolgt KI Forschung ab dem Jahr 2000 das Ziel, die Kommunikation zwischen zwei Systemen zu verbessern. Die Frage lautet stets: wie kommunizieren zwei Menschen miteinander? Wie interagiert ein Mensch mit einem Roboter? Welche Interaktionsspiele zwischen einem Speaker und einem Hearer gibt es?
Selbstverständlich ist die obige Beschreibung stark vereinfachend. Auch vor dem Jahr 2000 gab es Versuche die Mensch-Maschine Interaktion zu verbessern:
- 1968,SHRDLU natural language understanding by Terry Winograd
- 1987,Vitra visual translator
- 1998,Rocco Robocup commentator by Dirk Voelz
Man könnte z.B. SHRDLU als Vorläufer heutiger Vision Language action (VLA) Modelle beschreiben. Anzahlmäßig waren jedoch KI Projekte, die natürliche Sprache als Kommunikationsschnittstelle verwendeten eher die Ausnahme. Es war vor dem Jahr 2000 unklar, dass dies wichtig ist und es war unklar wie man eine solche Schnittstelle technisch realisieren könnte. Hinzu kommt dass ein interdisziplinärer Ansatz bestehend aus Mathematik und Lingustik vor dem Jahr 2000 unüblich war.
July 28, 2026
Graph traversal with a head up display
The perhaps most simple example for a head up display is a graph traversal problem of a robot. The robot moves inside a graph and should reach a target node.
The AI for the robot works with a head up display. There is a text box at the bottom showing the inner voice of the robot. The inner voice determines at which position the robot is, which nodes are in the near, what the target node is, and which action should be taken next.
A mathematical problem, graph traversal, gets converted into a textual problem. Textual means, that the head up display is using words to describe the reality. possible words are [currentnode, goalnode, nextnode, distance_to_goal]. These words and events are used to describe the game state from a high level perspective. The text box ensures that the inner voice was implemented correctly. That means, the AI isn't solving an optimization problem and its not running an algorithm, but the main task for the AI is to generate textual output in the head up display and talk to the human operator.
July 27, 2026
The paradigm shift in robotics around the year 2010
According to published research papers, the year 2010 was a turning point in robotics research. After this year, higher effort was put into human to robot interaction with natural language. Projects from this time span were:
- M.I.T. forklift by Stefanie Tellex
- Marco route instruction following by Matt MacMahon
- Word2vec algorithm by Tomas Mikolov
These projects were started from around 2008 until 2013. Many smaller projects also tried to use natural language for robot control.
This development was different from robotics research until 2010. The years before this year, there was a search for sophisticated algorithms available like training algorithms for neural networks, path planning algorithms and SLAM algorithms. The search for novel algorithms was working with the same principle how common computer science is working. The idea was, that AI gets implemented on a computer, computers need an algorithm and the consequence is to develop dedicated robotics algorithms for solving tasks.
The problem with the algorithm centric perspective until 2010 was, that all these developed techniques were not powerful enough. Even advanced probabilistic path planning algorithms implemented on a multi core CPU are not able to control a warehouse robot. The problem is the reality gap. The robot assumes a different reality than the real reality and the algorithm can't bridge the gap. It makes no sense to program more additional software modules or create a larger database for the robot, but the principle of autonomous robotics in general has to be questioned.
This paradigm shift took place in published academic literature around the year 2010. Research papers written after this date put a higher emphasis on teleoperation and grounded language for robot control. The idea is, that the source of wisdom is located outside of the robot as a human operator and the task is to get access to this knowledge by asking the human in natural language.
Even if the principle sounds inaccurate, it can be scaled up towards more complex scenarios. In the easiest case the human to robot interface works with a list of predefined commands, in a more advanced setup a neural network can parse the instructions. In contrast to figure out algorithms for autonomous robots, the new paradigm is to build language parser and see a robot as an open system.
The paradigm shift around the year 2010 allows to create artificial intelligence. Nearly all the former problems in robot control can be solved with the open system paradigm. Its only a detail problem how to create a high level user interface, so that the human operator can provide general statements like "tidy up the kitchen" and the robot is doing the full task by its own.
Roughly spoken, robotics programmed after the year 2010 are entirely teleoperated. There is a human operator in the background who gives instructions, or the former human operator instruction list was translated into a computer program who talks with the robot as large language model. in all the cases the robot interacts with a higher instance outiside of the robot and the artificial intelligence is located in the language interface.
July 25, 2026
Engine for grounded language
One possible explanation why the symbol grounding problem has emerged late in the history of computer science is because the theory is difficult to realize on a computer. Suppose natural language is important for robot control, the problem is create a language parser which is working for a concrete domain.
A possible command for a robot might be "Move until obstacle and then stop". Each of the words is stored as a string, but it remains unclear who to process the instruction into actions for a robot. The reason is, that the sentence is formulated in English but computers need a programming language as input. Even if every word is encoded as a number, it doesn't make sense to submit an array with numbers to the robot because its not possible to add or subtract the values in a meaningful way.
In general the problem is how to convert natural language into a computer program. Without solving this issue, the symbol grounding problem remains only a philosophical problem without any practical consequences.
The good news is, that the problem of programming a parser can be solved. Not with tools from computer science but by using techniques from linguistics, namely language games. Instead of treating language parsing as an algorithm problem, the idea is to invent around words a puzzle game. Typical language games are:
- Name guessing game. Player1 points to an object in the reality, and Player2 has to tell the name
- NPC quest game, a non player character in a role playing game formulates a quest like "bring me the sword from the wood" and the player has to fulfill the task
- bounding box game, player1 says a word like "table" and player2 has to draw a bounding box around this object
All these language games are located outside of computer science. They have nothing to do with algorithms, programming language nor existing robotics libraries, but they are games played with 2 human players.
The interesting situation is, that its possible to simulate the games with a computer. The software encodes the rules of the language game, determines the score for the human player and then the player can take action inside this game.
The problem is not how to program a certain parser, but the problem is how to formulate the game outside of a computer first. A well formulated game can be implemented in a software with ease. The programmer needs only the specification of the game including its rule, and then its possible to program the game with python. The only requirement is, that the computer works like the original language game. The programming workflow is identical to implement card games and board games on a computer.
July 24, 2026
Die späte Entdeckung der Schrift im Kontext von Robotik
Die Schrift ist eine sehr alte Erfindung der Menschheit. Die erste Bilderschrift, die Ägyptischen Hieroglyphen entstanden um 3200 v. Chr. Insgesamt ist Schrift und natürliche Sprache sehr detailiert erforscht. Es gibt umfassende Wörterbücher, Darstellungen welche die Geschichte der Sprache zeigen und Untersuchungen bezüglich Wortherkunft.
Vereinfacht gesagt sind Wörter Referenzsysteme zur Realität. Substantive stehen für Objekte wie "Himmel, Tisch, Apfel", Adjektive stehen für Tätigkeiten wie "Laufen, springen, geben" und Adjektive werden als Eigenschaftswörter verwendet wie "gelb, groß, schnell, feucht". Das Wissen bezüglich Wortarten und die Nennung von Beispielwörtern ist banal, allerdings nur für die Sprachwissenschaft selber. Im Bereich Computerwissenschaft und Mathematik wurde natürliche Sprache lange Zeit ignoriert. Es gab zwar Versuche chatbots zu programmieren, aber das war nur ein Teilbereich der Künstlichen Intelligenz.
Erst in jüngerer Zeit stellte sich heraus, dass natürliche Sprache womöglich das fehlende Puzzleteil darstellt mit der man künstliche Intelligenz inbesamt realisieren kann. Und zwar indem man Sprache als Technologie verwendet. Insbesondere dessen Eigenschaft auf die Realität zu verweisen macht es zum idealen Abstraktionsmechanismus. Es müssen keine neuen Sprachen erfunden werden sondern vorhandene Sprachen wie English, Deutsch usw. bieten bereits ein umfassendes Vokabular was von Robotern ähnlich wie Menschen verwendet werden kann. Alles was eine Maschine dafür benötigt ist eine Übersetzungstabelle von Bildern zu Sprache und in umgekehrter Richtung. Mit Hilfe dieser Bild zu Wort Tabelle kann man einem Roboter ein Kommando geben wie "Fahre zum Tisch". Und der Roboter übersetzt den Satz dann in Bilder und in Aktionen.
Bis ungefähr zum Jahr 2000 hat die KI Forschung nach Algorithmen gesucht mit deren Hilfe sich denkende Maschinen konstruieren lassen. Typische Algorithmen waren Lernverfahren für neuronale Netze, SLAM Algorithmen zur Selbstlokalisierung, A* Pfadplanungsalgorithmen oder Momdel predictive control Algorithmen. Die Annahme lautete jeweils dass mit diesen Algorithmen denkende Maschinen konstruiert werden könnten. Diese Annahme ist jedoch falsch. Es liegt zusätzlich der Verdacht nahe, dass es generell keine Algorithmen gibt, die Künstliche Intelligenz erzeugen, weil ein Algorithmus per se nicht mächtig genug ist um Roboter zu steuern. Was man stattdessen verwenden könnte wäre natürliche Sprache als zentrales Koordinierungsinstrument. Sprache ist ein Interface zwischen den Wortsymbolen einerseits und der Realität andererseits. Dadurch kann die komplexe Realität in wenige Wörter komprimiert werden.
Erst durch diese Realitätskompression ist es möglich den Handlungsraum für Roboter zu verkleinern. Der Roboter plant nicht länger in einem 3d Raum mit Trajektorien sondern der Roboter plant mit Hilfe von Verben und Substantiven in einem abstrakten Sprachraum. Seit der Erfindung von Word embeddings wie Word2vec mag dieser Ansatz selbstverständlich klingen aber bis zum Jahr 2000 war der Fokus auf Sprache eine Revolution.
Noch immer steht natürliche Sprache ein wenig außerhalb der klassischen Computerwissenschaft. Es hat nichts zu tun mit Elektrotechnik, Mathematik oder Algorithmen, sondern die Sprachwissenschaft hat ihren Ursprung in den Geisteswissenschaften. Nicht nur in der Dewey Dezimalklassifikation wie sie in Bibliotheken verwendet wird, sondern auch in der Gliederung von Universitäten sind Geistes- und Naturwissenschaften unversöhnliche Gegensätze die getrennt betrachtet werden.
Open systems for robotics
Robotics in the past was organized with a closed system paradigm. A robot was described as a machine which consists of hardware, software and algorithms and the task for the programmer was to improve the internal mechanism of the robot. It was ignored that robots are communicating with the outside world. For example a robot might receives commmands by teleoperation and submits a status code to the operator. Such kind of interaction was mostly described as wrong path towards robotics because such a machine isn't autonomous anymore. There decision making isn't determined by the internal algorithm but from the outside which was seen as anti pattern in Artificial intelligence.
It takes decades until computer science has questioned the self created bias. Modern robotics is working as open system which means, that the robot gets information from sensors and from remote control. Also the robot interacts with human operators and is able to answer questions like "What object is visible in the camera?".
The transition from closed to open systems in robotics can be seen as an important innovation. In contrast to invent yet another path planning algorithm or program a robot control software in C/C++ the open system paradigm reformulates the goals of a robot system. It puts a higher importance on the robot's environment and allows the enviornment to take influence on the robot. There are many examples available in the history of robotics with this background, e.g. Braitenberg vehicle, kismet social robot and SHRDLU. These projects have demonstrated interactive robotics. There is always a robot and a human operator who interacts with the robot.
From a technical perspectives, interactive robotics is equal to teleoperation. Teleoperation was recognized by computer science as opposite to artificial intelligence, because the machine doesn't decide by itself but is guided by external human wisdom. So the maschine can't be called a robot anymore but has more in common with a RC Car.
The rejection of teleoperation makes sense on the first look. If a human operator is in charge to control the RC car, then no artificial intelligence is needed. Therefor it has nothing to do with thinking machines and is located outside of robotics. Only autonomous robots are intelligent robots.
With a modern perspective, Artificial intelligence isn't located inside of a robot but its the interface between a robot and its environment. Such an interface can become smart in the sense that the interface understands natural language.
Practical demonstration of the Total turing test
Stevan Harnad coined around the year 1990 the term "total turing test" which is a philosophical description of an instruction following task in robotics. What is missing is a practical demonstration of such a test for a real robot.
Such a demonstration can be realized in a simple video game modeled as a language game. An entry level example is a navigation task in a graph. There are 8 nodes connected with lines and the robot has to move along the graph to reach a certain goal node. Possible interaction with the robot would be:
- what is your position?
- what is your battery status?
- which nodes are reachable from current position?
- Move north
- move to node #3.
- what is the shortest path to reach node #6?
The robot is in charge to answer these requests in natural language and with motor actions. The problem is easy enough to get implemented as normal computer code without using advanced large language models or vision language action models. The human to robot interaction can be simplified by using a codebook. The amount of possible commands is given in the menu and the human can select one of the commands. IN other words, the Total Turing Test (TTT) is some sort of speaker to hearer interaction game played between a human and a robot.
July 22, 2026
Grounded language in open systems
The box on the left is the human who describes the reality with natural language. The box on the right is the environment which can be perceived with sensors. Symbol grounding is the connection between both boxes.
From a system perspective the 2 box system is an open system because both boxes are connected to each other. Natural language from the left box is referencing to physical objects in the right box, while perceived reality in the right box gets described with English words in the left box.
The assumption is, that there are 2 different systems available which are working with different internal logic. The language layer consists of nouns, verbs, adjectives and grammars which is the symbolic layer. In contrast, the environment has no natural language but it consists of sensor perception, motor actions and 3d objects. The 2 box paradigm describes in a simplified format what natural language is about. Its an abstraction mechanism for the reality. Physical objects like a table or a banana are labeled with words. The ability to label objects is the key element in grounded language and allows to build intelligent robots.
Natural Language as AI technology
Languages likes English or French are discussed by Linguists not by computer scientists. A language is located in the humanities but not within the mathematics department. So its logical that most Linguists have no idea about computer science and vice versa. This might explain why AI research wasn't succesful over decades, because Natural langauge is the missing puzzle piece to make machines intelligent.
From a birds eye perspective, a language like English consists of verbs, nouns and adjectives. The grammar consists of rules how to connect the words to sentences. And language also is directed towards the reality. A word like "a flying bird" or "green flower" is referencing to objects which are available in front of the speaker.
The ability to name every object from the reality with a word and describe activities also with words makes natural language a powerful technology which can be used for human to human communication and machine to human communication both. In case of Artificial Intelligence a computer needs to parse natural language which is the bottleneck in modern AI research. Suppose a computer program asks the user to enter a word. The user enters "red box", then the computer stores the input in a variable but it has no consequences. So the computer isn't able to understand the meaning.
This parsing problem can be solved by inventing a language game. A language game is similar to a 2d arcade game a rule book which explain who to react to a certain input pattern. Typical language oriented games are translation games, the board game Scrabble or the "guess what" game. After implementing these games on a computer, the input of a human will have consequences. The consequences are given by the rules of a certain game, for example in a translation game the user needs to enter the correct translation for a word from another language. The computer verifies if the answer is correct.
Strictly spoken, not the computer decides about the meaning but the rules of the language games are providing the meaning. Modern AI research after the year 2020 is mostly focused on natural language and language games like "Visual question answering", "instruction following" and "question answering chatbots". All these games are formalizing the human to machine interaction in a sense that the computer can determine a score. This score allows to train artificial neural networks. The result of the training process is recognized as Artificial Intelligence in the modern sense. It allows to control robots with language instructions, and generate text with Large language models.
A common assumption in the past was that its very complicated to parse natural language with a computer. This assumption is only correct if the entire corpus of English should be understand by a computer which is around 1 million words and stored in endless amount of books. Parsing this written language with a computer is indeed a hard problem for computer science. What is possible instead is to reduce the task to a subset of English which consists of a dozens words from a restricted domain which are used to play a language game. In the minimal case, there are 6 picture cards and 6 word cards and the task is to match the correct pairs. Such a language game can be implemented in a short computer program and can be played by an automated AI algorithm. The AI algorithm has access to a database with the correct answers. This allows the computer to connect the picture of a banana with the word "banana".
July 21, 2026
Head up display for a kitchen robot
The picture shows an artist version of a head up display. It contains of:
- camera picture of a kitchen
- text box with inner voice
- bounding boxes
- labels for the bounding boxes
Surprisingly, the information in the picture can solve the symbol grounding problem because the head up display connects visual perception with textual information. The text from the inner voice like "I need to find 200g of flour" can be converted into meaning with the help of the bounding boxes. There is a box available with such an ingredient. The task for the robot is not to plan actions but the main problem is to connect language from the inner voice with detected objects in the camera.
Such a link of visual objects with textual labels is the core element in grounded language. If the robot is able to identify objects from the text box, its possible to generate all sort of inner voice. For example, the robot can say that he needs to peel the banana or "open the oven". All these nouns and verbs are translated into position of the bounding box on the screen which allows to execute the action physically.
July 20, 2026
How important is mathematics to understand Artificial intelligence?
On the first look, computer science has its root in mathematics and physics, therefor the assumption is, that AI is based on mathematics too. The precondition, according to the claim, to understand modern robotics is located in analysis, algebra, statistics and boolean algebra.
A closer look will show that mathematics isn't needed in AI research at all or to be more precise, mathematics didn't enable advanced AI in the past. What researchers from 1960 until around 2000s tried was to describe artifciial intelligence in terms of mathematics and logic and they failed. The problem is that none of the mentioned disciplines like statistics, algebra and so on contain a method how to enable artificial intelligence. Even if mathematics is a great science disciplines its complete useless to control a robot. Some attempts were made to solve robotics problems with mathematical description of trajectories including model predictive control, but even these advanced math subjects are not powerful enough to enable autonomous robots.
Artificial intelligence as a science discipline is working different than classic sciences like physics, math and computer science. From a pessimistic perspective there is no such thing like Artificial intelligence. Most research around thinking machines and autonomous robots comes to the conclusion that the promising new technique isn't working. That means, a large scale robot project with hundreds of man year effort isn't working at all, the applied techniques are useless and the researchers have no idea about the cause.
Such a kind of pessimistic situation is not an exception in AI research but it was the default situation from 1960 to 2000s. So the best analogy is to compare AI with a complex puzzle which is impossible to solve and no matter how well the researchers are familiar with mathematics, philosophy or psychology they had no idea how to make a machine intelligent.
AI research in the past was mostly a trial by error meta disciplines which wasn't able to solve any of the goals. It was impossible to build intelligent robots, it was hard to program AI agents for computer games, and speech recognition with a computer was another unsolved topic.
The good news is, that its possible to list techniques who are not leading to AI. These techniques are:
- neural networks
- maschine learning
- reinforcement learning
- mathematical optimiziation
- heuristics
- number crunching
- genetic algorithms
- state space search
- case based reasoning
- expert system
In other words, the entire AIMA book (Russle/Norvig: AI a modern approach) is an anti pattern. It doesn't contain recipes to program robots but it describes the struggle of AI researchers for doing so.
After listing all these techniques who are not leading to AI there is a need to find the shared bias. THis bias is the computational paradigm which means, to use a computer to provide intelligence. This shared bias hasn't worked in the past because its an anti pattern in AI research.
On the first look it makes sense to assume, that AI has to do with computation because the computer or the robot should calculate something which leads to intelligent decision making. Therefor AI has to do with number crunching, programming and mathematics. The problem is that in the reality this paradigm isn't able to control robots but it blocks the progress in technology.
It seems, that AI aka intelligence isn't located inside a robot but outside of the machine. On the first look, such an assumption sounds like blasphemy because outside of a computer there is nothing which can calculate or make decisions. At least until the 2000s such a claim would be rejected by mainstream AI research for sure. With more recent understanding of AI there are some new results available which show that AI might located indeed outside of a robot.
Suppose the source of intelligence is not located in the CPU and not inside of a robot hardware, then intelligence has nothing to do with mathematics or physics. This doesn't mean that intelligence is equal to a magic force but it implies that AI has to do with communication. Communication is the science of how to connect things, communication puts a focus on the air gap between two systems.
The transition from former computational paradigm which locates AI Inside of a robot, towards modern communication paradigm which locates AI between two systems is the major paradigm shift in AI research which took place after the year 2000. The revised understanding of intelligence is strongly connected with communication, linguistics and man to machine interaction. In contrast, former focus on computation, mathematics, algorithms and programming have been discarded.
July 19, 2026
Robot control with head up displays
In contrast to a famous assumption, modern robotics isn't working with algorithms or neural networks but the basic building block is graphical user interface, namely a head up display (HuD). The HuD solves the symbol grounding problem. Typical elements are: bounding boxes around detected objects, text labels for describing the content of a bounding box, another text box for showing the inner voice of rhe robot.
These ingredients are enough to program an advanced artificial intelligence which can solve complex problems. The HuD including the mentioned bounding boxes acts as a communication layer. It ensures that the computer understands basic commands like "move to shelf and grasp the box". A certain high level command is converted into a visual pictures in the HuD, e.g. the word "shelf" is referencing to a bounding box with the label "shelf" which has a 2d position on the screen.
Programming a Head up display for an existing video game is a demanding task but can be solved with standard programming techniques. Most videogames created since the 1980s have a built in debug mode which comes close to a head up display. In the debut mode, all the sprites on the screen are highlighted with frames and sometimes the name of the objects are shown as textual overlay. The combination of graphical display plus textual overlay is the main principle of a head up display and also the main principle of grounded language. So the HuD itself acts as technology for enabling artificial intelligence.
Let me give another example to demonstrate the advantages: Suppose the head up display for a warehouse robot videogame was activated. The user sees some bounding boxes on the screen for highlighting objects in the map like charging station, corridor, shelf A, shelf B, green box, red box. Also the inner voice of the robot is shown a text frame and contains:
"I'm standing at position (3,2). My battery level is 80%, my goal is to fetch the red box from shelf A, the planned trajectory is shown as arrows in the map".
So the initial situation for the robot is, that an annotated HuD is visible which labels objects and mentions the current goal. These information can be translated into actions for the robot. All what the AI of the robot has to do is to compile these information and decide what to do next. From an AI perspective its an instruction following task with an aciivated head up display.
A head up display provides a cognitive space. The shown bouding boxes and labels are creating a symbolic representation of the world. The world of the robot can be described in terms from the head up display. Its no longer a mathematical space and not a 3d space but the reality introduced by the HuD consists of words, locations of items and goals from the inner voice. Such a high level space can be processed by a computer because the amount of possible states is small. There are not millions of possible objects but the HuD shows only 6 different objects in a map. and the inner voice doesn't display millions of possible actions, but the inner voice describes clearly what the current situation is, and what the desired goal state is, similar to a text adventure.
July 18, 2026
Introduction into typst typeseetting
The amount of tutorials for typst is very low, because the software is new and works different from LaTeX. The following tutorial should explain the basics.
At first, a new file is created in the working directory which gets compiled into a pdf document with: "typst compile main.typ". The software itself is available as a binary file for all operating systems and needs around 60 MB on the SSD storage.
----------------
mainsimple.typ
----------------
#align(center)[
#text(24pt, weight:"bold", "title of paper")
#text(16pt, "Manuel Rodriguez\n 1 July 2026")
]
= 1 Introduction
#lorem(100)
= 2 Literature
- #lorem(10)
- #lorem(10)
If the typst user is reducing its demands to a minimum, the academic paper is ready for submission. Most authors have a need for more advanced layout so the file can be modified a bit.
----------------
maincomplex.typ
----------------
#set par(
justify: true,
spacing: 0.65em,
first-line-indent: 2em,
)
#set text(
font: "Liberation Sans",
size: 9pt,
lang: "en",
)
#align(center)[
#text(24pt, weight:"bold", "title of paper")
#text(16pt, "Manuel Rodriguez\n 1 July 2026")
]
#outline()
= 1 Introduction
#lorem(100)
#lorem(100)
= 2 Topic
== 2.1 Subtopic
#table(
columns: 2,
table.header[date][event],
[May 2, 2026], [hello],
[Jun 3, 2026], [world],
)
== 2.2 Subtopic
#figure(
image("drawing2.jpg", width: 4cm),
caption: [Drawing with pencil],
)
= 3 Literature
- #lorem(10)
- #lorem(10)
July 13, 2026
AI as open system
The last AI winter during the early 1990s was caused by the ignorance towards open systems namely Teleoperation for robot control. What the AI researchers have prefered instead were autonomous algorithm controlled AI systems. The goal was to program a large scale software similar to an operating system or a word processing software and make the software higly intelligent. Such bloat AI projects have failed, even a program written in 200k lines of code in C/C++ code isn't able to control a toy car in an obstacle course.
The reason why closed system failed is because existing programming languages like C/C++ can't grasp reality outside of a robot, existing algorithms like RRT pathplanning are too slow for realtime planning and the possible amount of trajectories for a robot in the reality is too large. This mixtures of challenges prevents that robot projects in the 1990s have become succesful. The written sourcecode was useless and the project doesn't make any sense.
The term AI winter is referencing to a situation in which the problems are known but no answer is available. This answer is maybe the transition from closed systems towards open systems. Open systems in robotics are equal to teleoperation which means that the human operator controls the robot. So its not longer an algorithm controlled robot but its an RC car. The main advantage is that such an open system can be realized easier with existing technology. The needed software is minimal and no true Artificial Intelligence is required.
What is used instead for remote controlled robots is a sender/receiver device which is an interface between human and machine. Such a device has multiple tasks:
1. it receives radio waves over the air
2. it parses natural language commands
3. it receives sensor signals
4. it transmits radio waves to the remove control
5. it transmits natural language status information to the human operator
In one word, the transceiver connects the robot with the environment.
The picture on the left shows the older paradigm. A robot in enclosed by a box and forms a closed system. The robot's AI is a turing machine executed on the CPU and the goal is to invent a sophisticated algorithm which makes the robot intelligent.
The picture on the right shows the modern paradigm which assumes two different systems connected by a sender/receiver. THese two systems are the robot and the environment around the robot. Both systems need to communicate. Communication doesn't require an autonomous algorithm but a protocol.
The transition from older closed systems into modern open systems is equal to discard algorithm oriented AI in favor of a communication perspective.
July 12, 2026
Die KI Blase ist geplatzt ...
Schauplatz: Ein Besprechungsraum am Rande der Endmontage in einem süddeutschen Automobilwerk.
Die Beteiligten:
Dr. Matthias Vogt (48), Leiter der Innovations- und Automatisierungsabteilung.
Elena Rostova (34), leitende Projektingenieurin für Robotik.
Dr. Julian Arndt (41), Key Account Manager von „Apex Robotics“ (Hersteller des Roboters).
Auf dem Tisch stehen drei unberührte Kaffeetassen. Durch die Glasscheibe sieht man die Werkshalle, in der ein leerer Stellplatz markiert ist. Die Testwoche des humanoiden Prototyps „Apex-One“ ist vorbei.
Arndt: (bemüht optimistisch) Erst einmal vielen Dank, Herr Dr. Vogt, Frau Rostova, dass wir unseren Apex-One unter echten Linienbedingungen testen durften. Ein Vision-Language-Action-Model, kurz VLAM, direkt in der Aggregate-Montage einzusetzen, das ist Pionierarbeit. Ich habe mir die Logdaten angesehen – die semantische Erfassung der Werkzeuge war phänomenal, oder nicht?
Vogt: (seufzt, reibt sich die Schläfen) Herr Arndt, ich mache es kurz. Der Roboter ist bereits verpackt. Er steht auf einer Palette im Wareneingang und wartet auf Ihren Spediteur. Wir treten von der Kaufoption zurück und werden das Projekt plangemäß beenden.
Arndt: (konsterniert) Bitte? Nach nur einer Woche? Gab es Hardware-Ausfälle? Wir können das Modell sofort gegen die Revision 1.4 austauschen, die hat verstärkte Aktuatoren in den Handgelenken…
Rostova: Es liegt nicht an den Gelenken, Herr Arndt. Es liegt am Gehirn. Genauer gesagt: an der Latenz und der mangelnden Deterministik dieses KI-Ansatzes.
Arndt: Aber das VLAM ist die Zukunft! Sie steuern die Maschine mit natürlicher Sprache. Keine Zeile Code. Der Roboter sieht die Werkstücke, versteht den Befehl und handelt.
Rostova: Ja, in der Theorie. In der Praxis sah das so aus: Am Dienstag sollte der Roboter Getriebeölkühler aus der Kiste nehmen und am Chassis fixieren. Der Befehl lautete: „Nimm den Kühler, überprüfe die Dichtung und setze ihn an Position B.“ Wissen Sie, was passiert ist?
Arndt: Er hat die Position gesucht?
Rostova: Er hat elf Sekunden lang „nachgedacht“. Elf Sekunden Standzeit, in denen sein neuronales Netz die visuelle Szene mit dem Sprachbefehl abgeglichen hat. In der Taktzeit unserer Produktion sind elf Sekunden eine Ewigkeit. Und als die Spätschicht am Mittwoch den Befehl leicht abwandelte – „Kühler greifen, Dichtring checken, ran an B“ – hat das Modell halluziniert. Es hat den Kühler gegriffen und ihn mit achtzig Newtonmetern gegen die Windschutzscheibe gedrückt, weil es „ran an B“ als „Scheibe einschlagen“ interpretiert hat.
Arndt: Oh. Gab es einen Personenschaden?
Vogt: Gott sei Dank nein, die Lichtgitter haben ausgelöst. Aber die Windschutzscheibe war Schrott und das Band stand für zwanzig Minuten. Herr Arndt, wir bauen hier achthundert Fahrzeuge am Tag. Wir können uns keine Maschine leisten, die auf denselben Befehl dreimal unterschiedlich reagiert, nur weil sich das Umgebungslicht ändert oder der Werker einen Dialekt spricht.
Arndt: Das sind Feinheiten im Prompt-Engineering! Wir können das Modell feintunen. Wir füttern es mit spezifischen Daten aus Ihrer Halle. Mit einem Ersatzmodell und zwei Wochen Datenkorrektur kriegen wir die Fehlerquote unter ein Prozent.
Vogt: Ein Prozent? Das ist im Automobilbau eine Katastrophe. Ein herkömmlicher Knickarmroboter von Kuka oder Fanuc arbeitet mit einer Wiederholgenauigkeit von weniger als einem Zehntel Millimeter, stundenlang, fehlerfrei, deterministisch. Er denkt nicht nach, er tut es einfach.
Arndt: Aber ein Knickarmroboter kann nicht flexibel auf unstrukturierte Kisten reagieren oder per Sprache umprogrammiert werden! Humanoiden sind die Zukunft für die flexible Montage.
Rostova: Flexibilität nützt uns nichts, wenn sie auf Kosten der Prozesssicherheit geht. Ihr Apex-One hat versucht, einen Schlagschrauber wie eine Kaffeetasse zu greifen, weil am Donnerstag jemand eine Mate-Flasche neben der Station vergessen hatte und das Vision-Modell die Geometrien verwechselt hat. Die Multimodalität ist für komplexe, dynamische Industrieanwendungen einfach noch nicht reif. Es fehlen die harten Sicherheitsgarantien. Ein neuronales Netz ist eine Blackbox. Wir können nicht zertifizieren, was wir nicht mathematisch beweisen können.
Arndt: (schaut auf seine Notizen) Ich verstehe Ihre Frustration. Aber bedenken Sie den Imagegewinn. Ein humanoider Roboter an der Linie…
Vogt: (unterbricht ihn kühl) …ist teures Theater für die Aktionärshauptversammlung, aber kein Werkzeug für die Werkshalle. Wir brauchen keine Roboter, die wie Menschen aussehen und versuchen, wie Menschen zu denken, nur um Aufgaben zu erledigen, die eine starre Automatisierungslösung in einem Zehntel der Zeit für ein Fünftel der Kosten erledigt.
Arndt: Also kein Ersatzmodell? Auch kein kostenloser Folgetest mit unserer neuesten Software-Generation im Herbst?
Vogt: Nein. Das Thema Humanoiden ist für uns vorerst gestorben. Wir investieren das Budget wieder in klassische Portalroboter und smarte Kamerasysteme. Die sprechen zwar nicht mit uns, aber sie halten den Takt.
Rostova: (steht auf) Ich begleite Sie zum Wareneingang, Herr Arndt. Die Papiere für die Rückgabe liegen beim Meister.
Arndt: (packt enttäuscht sein Tablet ein) Schade. Sie verpassen den Anschluss an die nächste industrielle Revolution.
Vogt: Mag sein. Aber dafür steht mein Band morgen früh um sechs nicht still. Auf Wiedersehen, Herr Arndt.
July 10, 2026
History of TeX from 1985-1995
For newbies in document typesetting, the current LaTeX ecosystem seems to be obsolete and populated with lots of useless packages. Its unclear about the the TeX community is talking exactly if they are discussing certain parameters for a certain LaTeX fork like Xelatex. To understand the current mess we have to go back some years into the past.
The dacade from 1985 until 1995 can be described as the rise of TeX. The system was using state of the art technology and made professional typesetting on a computer available for the mass. In 1985 Donald Knuth released Tex version 3.0 which evolved later into
3.141592653, also he invented the .dvi output format. In the year 1990 TeX become popular for a larger audience, due to the development of distributions which combined TeX, fonts, and additional programs, also the extension LaTeX was created in the early 1990s. Around the year 1995, LaTeX had become the standard in academic publishing. A .tex file was compiled into a postscript file including mathematical equations and postscript fonts which was revolutionary at this time.
Unfurtunately, the years after 1995 can be described as a decline in the TeX community. There are multiple problems available. First, Donald Knuth decided to freeze the development of the TeX engine, secondly lots of forks were created like latex3, omega, context, pdflatex, xetex and so on with the attempt to improve the original project. The CTAN archive was initially planned as a repository of useful packages, evolved into a messy museum of obosolete code. Instead of throwing away outdated code, font specification and templates, the TeX Community decided to preserve the past at any price.
10 years later around the year 2005, the LaTeX ecosystem showed the first sign of serious problems. The mainstream typesetting reality has switched to the pdf format and introduced HTML documents for the internet, while TeX users were devoted to the former dvi/postscript pipeline. It was very difficult to use foreign special characters and the amount of possible packages increased.
July 09, 2026
The slow emergence of Artificial Intelligence
AI and robotics was researched since decades. In contrast to other disciplines like computer science or mathematics there was no success available. Even if AI researchers have analyzed the subject from a scientific perspective and discussed the situation at conferences there were not able to identify major problems or offer possible answers. What was happen instead was a long disappointing journey.
Even if AI in the past suffered, lots of subjects were analyzed. Notable examples are: autonomous robotics, model predictive control, genetic algorithms, reinforcement learning, Turing maschines. All these subjects were seen as promising candidates towards the pathway to intelligent machinery. They can be called advanced subjects in computer science and many papers were written. The general idea was to describe intelligence as an optimization problem which can be measured with a score. For example, trajectory optimization tries to reduce the costs, while genetic algorithms are maximing the fitness of candidates. In both cases the computer is a device for solving a mathematical problem.
On the first look, it makes sense to describe robotics movement with model predictve control algorithms. It helps to translate a problem from the reality towards an abstract mathematical equation. The idea is, that artificial intelligence can be realized as a combination of computer science, mathematics and game theory. Most researchers in the past would agree, that such kind of interdisplinary approach is a sign of excellence and allows to discover future robotics algorithms. What the researchers in the 1990s and 2000s didn't know was that the describe workflow is a dead end. None of mentioned techniques lead to artificial intelligence.
Model predictive control is a good example for a dead end in robotics resarch. The subject was researched by multiple researchers independend from each other with a great effort. There is no obvious mistake in the equation nor in a certain paper about the subject. At the same time, the entire model predictive control research has to be called a dead end because it fails to control simple robots.
In the history of artificial intelligence such kind of dead end is not an exception but its default situation. All the other attempts to realize robotics like expert system, neural networks and 5th generation programming languages like Prolog have failed too. It seems that there was a need to explore all the non working principles to get a better understanding what sort of approach won't result into a working robot.
Ai research in the past was realized as an intersection of physics models, mathematical theories and computer science. The hope was that the combination of these powerful disciplines allows to create intelligent machinery. For example trajectory optimiziation has a background in theoretical physics, and can be implemented as an algorithm on a computer. This would allow to plan the movement of a robot.
What was unknown in the past or perhaps it was ignored was, that the state space in robotics is too large to use mathematical optimization problems. Predicting future states of a system is only possible if the system consists of a few variables e.g. in a predator-prey scenario modelled with Lotka–Volterra equations. Such a system can be calculated on a computer and future states can be processed in advanced. The concept fails on robotics domains like dexterous grasping or biped walking. The equations are not known or they are too complicated the calculate. Even if there are realistic physics simulators available like Box2D, its not possible to determine future states of these engines into the future.
Despite this pessimistic situation it makes sense to explore model predtctive control and other mathematical optimization techniques because it allows a better understanding of np hard problems. If its known, that the state space in robotics is too large, its possible to rethink about the situation and explore strategies how to reduce the state space. A state space reduction is the pathway to advanced robotics.
July 08, 2026
Microtype simulator in python
import pygame
import sys
import math
# Initialize Pygame
pygame.init()
pygame.font.init()
# Constants
WIDTH, HEIGHT = 1100, 780
SCREEN = pygame.display.set_mode((WIDTH, HEIGHT))
pygame.display.set_caption("Professional Microtype Engine & Layout Simulator")
CLOCK = pygame.time.Clock()
# Palettes
COLOR_BG = (249, 248, 245) # Premium archival paper
COLOR_TEXT = (35, 35, 35) # Soft black
COLOR_MARGIN = (230, 90, 90) # Margin guideline
COLOR_UI_BG = (225, 227, 230)
COLOR_UI_TEXT = (50, 55, 60)
COLOR_SLIDER = (70, 130, 180)
COLOR_ACTIVE = (46, 139, 87) # SeaGreen for scores/active selections
# Load fonts
try:
FONT_SIZE = 18
FONT = pygame.font.SysFont("georgia", FONT_SIZE)
FONT_BOLD = pygame.font.SysFont("georgia", FONT_SIZE, bold=True)
except:
FONT = pygame.font.Font(None, FONT_SIZE)
FONT_BOLD = pygame.font.Font(None, FONT_SIZE)
SAMPLE_TEXT = (
"Typography is the art and technique of arranging type to make written language "
"legible, readable, and appealing when displayed. The Knuth-Plass dynamic programming "
"algorithm revolutionizes this by looking ahead at the entire paragraph. Instead of "
"making hasty choices on a line-by-line basis, it distributes layout 'badness' evenly, "
"preventing unexpected blocks of loose text. Combined with microtype tracking expansions, "
"subtle margin protrusions yield pristine geometric columns resembling classic elite print."
)
# --- UI Widgets ---
class Slider:
def __init__(self, x, y, w, h, min_val, max_val, start_val, label):
self.rect = pygame.Rect(x, y, w, h)
self.min_val = min_val
self.max_val = max_val
self.val = start_val
self.label = label
self.grabbed = False
self.update_handle()
def update_handle(self):
ratio = (self.val - self.min_val) / (self.max_val - self.min_val)
hx = self.rect.x + int(ratio * self.rect.w)
self.handle_rect = pygame.Rect(hx - 5, self.rect.y - 4, 10, self.rect.h + 8)
def draw(self, screen):
lbl = FONT.render(f"{self.label}: {self.val:.2f}", True, COLOR_UI_TEXT)
screen.blit(lbl, (self.rect.x, self.rect.y - 22))
pygame.draw.rect(screen, (190, 195, 200), self.rect, border_radius=3)
pygame.draw.rect(screen, COLOR_SLIDER, self.handle_rect, border_radius=3)
def handle_event(self, event):
if event.type == pygame.MOUSEBUTTONDOWN:
if self.handle_rect.collidepoint(event.pos) or self.rect.collidepoint(event.pos):
self.grabbed = True
elif event.type == pygame.MOUSEBUTTONUP:
self.grabbed = False
elif event.type == pygame.MOUSEMOTION and self.grabbed:
mx = max(self.rect.x, min(event.pos[0], self.rect.x + self.rect.w))
rel = (mx - self.rect.x) / self.rect.w
self.val = self.min_val + rel * (self.max_val - self.min_val)
self.update_handle()
class RadioSelector:
def __init__(self, x, y, options):
self.x = x
self.y = y
self.options = options
self.selected_index = 1 # Default to Knuth-Plass
self.buttons = []
for idx, opt in enumerate(options):
bx = x + (idx * 280)
self.buttons.append(pygame.Rect(bx, y, 20, 20))
def draw(self, screen):
lbl_title = FONT_BOLD.render("Line Breaking Algorithm:", True, COLOR_UI_TEXT)
screen.blit(lbl_title, (self.x, self.y - 25))
for idx, opt in enumerate(self.options):
rect = self.buttons[idx]
# Draw outer circle
pygame.draw.circle(screen, COLOR_UI_TEXT, rect.center, 10, 2)
# Draw internal selection
if idx == self.selected_index:
pygame.draw.circle(screen, COLOR_ACTIVE, rect.center, 6)
lbl = FONT.render(opt, True, COLOR_UI_TEXT)
screen.blit(lbl, (rect.x + 25, rect.y + 1))
def handle_event(self, event):
if event.type == pygame.MOUSEBUTTONDOWN:
for idx, rect in enumerate(self.buttons):
# Expanded click zone for user convenience
click_zone = rect.inflate(150, 10)
if click_zone.collidepoint(event.pos):
self.selected_index = idx
return True
return False
# --- Helper Text Calculation Tools ---
def compute_word_widths(words, font, tracking):
return [sum(font.size(char)[0] + tracking for char in word) for word in words]
def calc_line_badness(width, test_width, num_gaps, base_space_width, min_space, max_space, ideal_space, is_last=False):
if num_gaps == 0:
remaining = width - test_width
return (remaining ** 2) if remaining >= 0 else 500000
actual_space = (width - test_width) / num_gaps
if actual_space < min_space:
# Heavily penalize over-compressed lines
return 100000 + (min_space - actual_space) * 50000
elif actual_space > max_space:
# Loose lines
return int(((actual_space - max_space) ** 2) * 500)
else:
# Standard deviation penalty
badness = int(((actual_space - ideal_space) ** 2) * 100)
if is_last and actual_space > ideal_space:
return 0 # Last line of a paragraph shouldn't stretch to fill the margin
return badness
def apply_protrusion(word, font, protrusion):
protruding_chars = [".", ",", "-", "!", "?"]
if protrusion > 0 and word[-1:] in protruding_chars:
return font.size(word[-1:])[0] * protrusion * 0.5
return 0
# --- Line-Breaking Core Algorithms ---
def layout_greedy(words, word_widths, font, width, min_space, max_space, ideal_space, protrusion):
""" a) Traditional First-Fit Greedy Algorithm """
lines = []
current_line, current_widths = [], []
current_width = 0
for idx, word in enumerate(words):
w_width = word_widths[idx]
p_adjust = apply_protrusion(word, font, protrusion)
# Test if it fits with standard spaces
test_w = current_width + w_width + (ideal_space if current_line else 0) - p_adjust
if test_w <= width or not current_line:
current_line.append(word)
current_widths.append(w_width)
current_width += w_width + (ideal_space if len(current_line) > 1 else 0)
else:
# Seal line
num_gaps = len(current_line) - 1
last_word_pad = apply_protrusion(current_line[-1], font, protrusion)
pure_width = sum(current_widths)
space_used = (width - (pure_width - last_word_pad)) / num_gaps if num_gaps > 0 else ideal_space
badness = calc_line_badness(width, pure_width - last_word_pad, num_gaps, ideal_space, min_space, max_space, ideal_space)
lines.append((current_line, current_widths, space_used, False, badness))
current_line, current_widths = [word], [w_width]
current_width = w_width
if current_line:
num_gaps = len(current_line) - 1
last_word_pad = apply_protrusion(current_line[-1], font, protrusion)
pure_width = sum(current_widths)
space_used = ideal_space
badness = calc_line_badness(width, pure_width - last_word_pad, num_gaps, ideal_space, min_space, max_space, ideal_space, is_last=True)
lines.append((current_line, current_widths, space_used, True, badness))
return lines
def layout_knuth_plass(words, word_widths, font, width, min_space, max_space, ideal_space, protrusion):
""" b) Look-Ahead Optimization (Global Minimum Variance) """
n = len(words)
dp = [(float('inf'), -1, ideal_space, 0)] * (n + 1)
dp[0] = (0, -1, ideal_space, 0)
for i in range(n):
if dp[i][0] == float('inf'): continue
current_width = 0
for j in range(i, n):
current_width += word_widths[j]
num_gaps = j - i
is_last = (j == n - 1)
p_adjust = apply_protrusion(words[j], font, protrusion)
line_txt_w = current_width - p_adjust
badness = calc_line_badness(width, line_txt_w, num_gaps, ideal_space, min_space, max_space, ideal_space, is_last)
actual_space = ideal_space
if num_gaps > 0 and not is_last:
actual_space = (width - line_txt_w) / num_gaps
p_cost = dp[i][0] + badness
if p_cost < dp[j + 1][0]:
dp[j + 1] = (p_cost, i, actual_space, badness)
lines, curr = [], n
while curr > 0:
parent = dp[curr][1]
if parent == -1: break
is_last = (curr == n)
lines.append((words[parent:curr], word_widths[parent:curr], dp[curr][2], is_last, dp[curr][3]))
curr = parent
lines.reverse()
return lines
def layout_first_fit_tight(words, word_widths, font, width, min_space, max_space, ideal_space, protrusion):
""" c) Alternating Minimum Space Greedy Algorithm """
# This variant forces as many words onto the line as physically allowed by compressing down to min_space limits.
lines = []
current_line, current_widths = [], []
for idx, word in enumerate(words):
w_width = word_widths[idx]
current_line.append(word)
current_widths.append(w_width)
p_adjust = apply_protrusion(word, font, protrusion)
num_gaps = len(current_line) - 1
min_needed = sum(current_widths) + (num_gaps * min_space) - p_adjust
if min_needed > width and num_gaps > 0:
# Overfilled line, dump the last token to the next row
popped_word = current_line.pop()
popped_width = current_widths.pop()
num_gaps = len(current_line) - 1
last_word_pad = apply_protrusion(current_line[-1], font, protrusion)
pure_width = sum(current_widths)
space_used = (width - (pure_width - last_word_pad)) / num_gaps if num_gaps > 0 else ideal_space
badness = calc_line_badness(width, pure_width - last_word_pad, num_gaps, ideal_space, min_space, max_space, ideal_space)
lines.append((current_line, current_widths, space_used, False, badness))
current_line, current_widths = [popped_word], [popped_width]
if current_line:
num_gaps = len(current_line) - 1
last_word_pad = apply_protrusion(current_line[-1], font, protrusion)
pure_width = sum(current_widths)
badness = calc_line_badness(width, pure_width - last_word_pad, num_gaps, ideal_space, min_space, max_space, ideal_space, is_last=True)
lines.append((current_line, current_widths, ideal_space, True, badness))
return lines
# --- Rendering ---
def render_paragraph(lines, font, x_start, y_start, tracking, protrusion, leading_ratio):
y = y_start
line_height = int(font.get_linesize() * leading_ratio)
for line_words, line_widths, space_width, is_last, _ in lines:
x = x_start
num_words = len(line_words)
for w_idx, word in enumerate(line_words):
for c_idx, char in enumerate(word):
char_surf = font.render(char, True, COLOR_TEXT)
render_x = x
if w_idx == num_words - 1 and c_idx == len(word) - 1:
render_x += apply_protrusion(word, font, protrusion)
SCREEN.blit(char_surf, (render_x, y))
x += char_surf.get_width() + tracking
if w_idx < num_words - 1:
x += space_width
y += line_height
# --- UI Layout ---
sliders = [
Slider(50, 540, 260, 10, -1.5, 3.0, 0.0, "Font Expansion (Tracking)"),
Slider(380, 540, 260, 10, 0.4, 1.0, 0.65, "Min Word Space Elasticity"),
Slider(710, 540, 260, 10, 1.0, 3.0, 1.70, "Max Word Space Elasticity"),
Slider(50, 620, 260, 10, 0.0, 1.2, 0.5, "Character Protrusion"),
Slider(380, 620, 260, 10, 0.8, 2.5, 1.3, "Line Height (Leading)")
]
algo_radio = RadioSelector(50, 710, ["a) Greedy Algorithm", "b) Knuth-Plass Ahead", "c) Space-Tight Fit"])
MARGIN_LEFT = 200
BOX_WIDTH = 700
# Main loop
while True:
SCREEN.fill(COLOR_BG)
# Event Engine Loop
for event in pygame.event.get():
if event.type == pygame.QUIT:
pygame.quit()
sys.exit()
for slider in sliders:
slider.handle_event(event)
algo_radio.handle_event(event)
# Drawing background infrastructure boundaries
pygame.draw.rect(SCREEN, COLOR_UI_BG, (0, 480, WIDTH, HEIGHT - 480))
pygame.draw.line(SCREEN, (190, 195, 200), (0, 480), (WIDTH, 480), 2)
pygame.draw.line(SCREEN, COLOR_MARGIN, (MARGIN_LEFT, 75), (MARGIN_LEFT, 450), 1)
pygame.draw.line(SCREEN, COLOR_MARGIN, (MARGIN_LEFT + BOX_WIDTH, 75), (MARGIN_LEFT + BOX_WIDTH, 450), 1)
# Gather metrics
base_space_width = FONT.size(" ")[0]
tracking_val = sliders[0].val
min_space = base_space_width * sliders[1].val
max_space = base_space_width * sliders[2].val
protrusion_val = sliders[3].val
leading_val = sliders[4].val
# Re-tokenize and check widths inside runtime
words = SAMPLE_TEXT.split(" ")
word_widths = compute_word_widths(words, FONT, tracking_val)
# Route processing via radio flag selections
if algo_radio.selected_index == 0:
computed_lines = layout_greedy(words, word_widths, FONT, BOX_WIDTH, min_space, max_space, base_space_width, protrusion_val)
elif algo_radio.selected_index == 1:
computed_lines = layout_knuth_plass(words, word_widths, FONT, BOX_WIDTH, min_space, max_space, base_space_width, protrusion_val)
else:
computed_lines = layout_first_fit_tight(words, word_widths, FONT, BOX_WIDTH, min_space, max_space, base_space_width, protrusion_val)
# Cumulative Badness Score Calculation
total_paragraph_badness = sum(line[4] for line in computed_lines)
# Render Paragraph Blocks
render_paragraph(computed_lines, FONT, MARGIN_LEFT, 95, tracking_val, protrusion_val, leading_val)
# Render Widgets
for slider in sliders:
slider.draw(SCREEN)
algo_radio.draw(SCREEN)
# Display Badness score at the top panel
score_lbl = FONT_BOLD.render(f"Overall Paragraph Badness Score: {total_paragraph_badness}", True, COLOR_ACTIVE)
SCREEN.blit(score_lbl, (MARGIN_LEFT, 35))
pygame.display.flip()
CLOCK.tick(30)






