Showing posts with label Grounding problem. Show all posts
Showing posts with label Grounding problem. Show all posts

September 29, 2026

Text based pointing game for car driving

One problem with grounded language that its hard to give a sense making example which can be implemented in a short amount of code lines. A possible answer is a NPC quest generator which is generating tasks as textual output. The following python program has only 33 lines of code and generates random quests for a car driving game.

All the quests are pointing challenges, the human is asked to point at a certain object in the scene. The software can't verified if the human has pointed at the correct object, but only the quest itself is shown on the screen.

The implementation in python is as simple as possible. There is a python dict with possible target location which are selected randomly by the software. So its some sort dictionary with a random element. To solve the tasks, that human needs to know what a certain word means. For example "Focus on the pedestrian on sidewalk." is asking for a certain object at a certain position. 

screenshot:

==================================================
  3D DRIVING GAME: NPC POINTING QUEST GENERATOR
==================================================
Press [ENTER] for a new quest | Type 'q' + [ENTER] to quit

[Quest #1] Ready? 
 >> NEW QUEST: Focus on the pedestrian on sidewalk.
--------------------------------------------------
[Quest #2] Ready? 
 >> NEW QUEST: Focus on the parking spot.
--------------------------------------------------
[Quest #3] Ready? 
 >> NEW QUEST: Track the parking spot.
--------------------------------------------------
[Quest #4] Ready? 
 >> NEW QUEST: Locate the car on the left lane.
--------------------------------------------------
[Quest #5] Ready? 
 

source code:

import random

# Driving game targets organized by relative 3D perspective from behind the wheel
TARGETS = {
    "traffic_controls": ["traffic light", "speed limit sign", "stop sign", "yield sign"],
    "road_features": ["street ahead", "crosswalk", "lane line", "pothole", "guardrail"],
    "vehicles": ["car in front", "car on the left lane", "car on the oncoming side", "truck in side mirror", "motorcycle in rearview mirror"],
    "environment": ["pedestrian on sidewalk", "billboard", "parking spot", "street lamp"]
}

VERBS = ["point at", "focus on", "locate", "target", "track"]

def generate_quest():
    category = random.choice(list(TARGETS.keys()))
    obj = random.choice(TARGETS[category])
    verb = random.choice(VERBS)
    return f'{verb.capitalize()} the {obj}.'

def main():
    print("=" * 50 + "\n  3D DRIVING GAME: NPC POINTING QUEST GENERATOR\n" + "=" * 50)
    print("Press [ENTER] for a new quest | Type 'q' + [ENTER] to quit\n")
    
    count = 1
    while True:
        cmd = input(f"[Quest #{count}] Ready? ").strip().lower()
        if cmd == 'q':
            print("\nGenerator stopped. Keep your eyes on the road!")
            break
        print(f" >> NEW QUEST: {generate_quest()}\n" + "-" * 50)
        count += 1

if __name__ == "__main__":
    main()


September 28, 2026

Pointing game with npc quests

 The game engine generates a random quests like "click on top" or "click on obstacle". The human user has to follow the instruction to get a reward.

import sys, random, pygame

pygame.init()
pygame.font.init()

# Setup grid & display settings
GRID_SIZE, CELL_SIZE = 8, 60
OFF_X, OFF_Y = 160, 60
W, H = 800, 650

SCREEN = pygame.display.set_mode((W, H))
pygame.display.set_caption("Grounded Language: Grid Pointing Game")
CLOCK = pygame.time.Clock()

FONT_MAIN = pygame.font.SysFont("Arial", 20, bold=True)
FONT_UI = pygame.font.SysFont("Arial", 16)

# Colors: BG, Grid, Obstacle, Box BG, Box Border, Text, Success, Error
C_BG, C_GRID, C_OBS = (240, 243, 246), (180, 185, 190), (100, 110, 120)
C_BOX_BG, C_BOX_BDR, C_TXT = (220, 225, 230), (140, 150, 160), (30, 40, 50)
C_SUCC, C_ERR = (46, 204, 113), (231, 76, 60)

# 2x2 Obstacle cells in the middle
OBSTACLE_CELLS = {(3, 3), (3, 4), (4, 3), (4, 4)}

ACTIONS = ["point at", "click on", "select", "target"]

class PointingGame:
    def __init__(self):
        self.score = 0
        self.fb_msg, self.fb_col, self.fb_timer = "", C_TXT, 0
        self.new_quest()

    def new_quest(self):
        act = random.choice(ACTIONS)
        qtype = random.choice(["specific", "obstacle", "top", "bottom", "left", "right"])

        if qtype == "specific":
            x, y = random.randint(0, 7), random.randint(0, 7)
            self.quest = f'NPC Quest: "{act} ({x},{y})"'
            self.check = lambda cx, cy, x=x, y=y: (cx, cy) == (x, y)
        elif qtype == "obstacle":
            self.quest = f'NPC Quest: "{act} the obstacle"'
            self.check = lambda cx, cy: (cx, cy) in OBSTACLE_CELLS
        else: # Top, Bottom, Left, or Right edge (8 cells each)
            self.quest = f'NPC Quest: "{act} {qtype} edge cells"'
            edges = {
                "top": lambda cx, cy: cy == 0,
                "bottom": lambda cx, cy: cy == 7,
                "left": lambda cx, cy: cx == 0,
                "right": lambda cx, cy: cx == 7
            }
            self.check = edges[qtype]

    def handle_click(self, pos):
        mx, my = pos
        if OFF_X <= mx < OFF_X + 480 and OFF_Y <= my < OFF_Y + 480:
            cx, cy = (mx - OFF_X) // CELL_SIZE, (my - OFF_Y) // CELL_SIZE
            if self.check(cx, cy):
                self.score += 10
                self.fb_msg, self.fb_col = f"Correct! +10 pts for cell ({cx},{cy})", C_SUCC
                self.new_quest()
            else:
                self.score = max(0, self.score - 5)
                self.fb_msg, self.fb_col = f"Wrong! Cell ({cx},{cy}) is not the target. (-5)", C_ERR
            self.fb_timer = pygame.time.get_ticks() + 1500

    def draw(self):
        SCREEN.fill(C_BG)
        # Headers & Score
        SCREEN.blit(FONT_MAIN.render("Grounded Language Grid Game", True, C_TXT), (OFF_X, 15))
        SCREEN.blit(FONT_MAIN.render(f"Score: {self.score}", True, C_SUCC), (OFF_X + 320, 15))

        # Axis Labels & Grid
        mx, my = pygame.mouse.get_pos()
        for i in range(8):
            SCREEN.blit(FONT_UI.render(str(i), True, C_TXT), (OFF_X + i * 60 + 26, OFF_Y - 25))
            SCREEN.blit(FONT_UI.render(str(i), True, C_TXT), (OFF_X - 25, OFF_Y + i * 60 + 22))

        for cy in range(8):
            for cx in range(8):
                rect = pygame.Rect(OFF_X + cx * 60, OFF_Y + cy * 60, 60, 60)
                col = C_OBS if (cx, cy) in OBSTACLE_CELLS else (255, 255, 255)
                pygame.draw.rect(SCREEN, col, rect)
                pygame.draw.rect(SCREEN, C_GRID, rect, 1)

                if rect.collidepoint(mx, my):
                    s = pygame.Surface((60, 60), pygame.SRCALPHA)
                    s.fill((52, 152, 219, 100))
                    SCREEN.blit(s, rect.topleft)

        # Quest & Feedback Widget
        box = pygame.Rect(OFF_X - 40, OFF_Y + 505, 560, 85)
        pygame.draw.rect(SCREEN, C_BOX_BG, box, border_radius=8)
        pygame.draw.rect(SCREEN, C_BOX_BDR, box, width=2, border_radius=8)

        SCREEN.blit(FONT_MAIN.render(self.quest, True, C_TXT), (box.x + 20, box.y + 15))
        if pygame.time.get_ticks() < self.fb_timer:
            SCREEN.blit(FONT_UI.render(self.fb_msg, True, self.fb_col), (box.x + 20, box.y + 50))

game = PointingGame()
while True:
    for event in pygame.event.get():
        if event.type == pygame.QUIT:
            pygame.quit(); sys.exit()
        elif event.type == pygame.MOUSEBUTTONDOWN and event.button == 1:
            game.handle_click(event.pos)

    game.draw()
    pygame.display.flip()
    CLOCK.tick(60)
 

September 25, 2026

Programming grounded language games step by step

Suppose the goal is to use grounded language to control a model railroad. The first step is to invent a dictionary with useful words:

- locomotive_1: red, locomotive_2: blue
- switch_1: left_side, switch_2: right_side, track,
- slowdown, speedup, switch, follow, wait

These words are stored in a python dictionary as a list. The first iteration of the parser takes a command from the command line e.g. "switch_1" and searches in the dictionary if the found is available. Then the parser returns "ok".

In step 2 the vocabulary gets connected with the visual appearance of the game formalized in a pointing game. The user enters a command and the parser should highlight the object on the screen. For example the user enters "locomotive_2" and the parser draws a rectangle around the object.

In step 3 which is more advanced an instruction following game gets established. The parser has to execute actions. The user might enter a command like "speedup locomotive_1" and the parser ensures that the desired action gets executed. Programming such a behavior is the most advanced part of a language game.

In general grounded language starts always with a vocabulary which is a word list. The list contains of nouns, verbs and objects and is related to a domain. Symbol grounding in the strict sense means to play language games with the word list which are the pointing game and the instruction following game.

Wo kann man grounded Language einordnen?

 Wissenschaft ist in Gebieten organisiert wie Mathematik, Physik, Linguistik und Kunst. Leider ist es schwierig, die Thematik "grounded language" einem dieser Bereiche zuzuordnen. Gleichzeitig ist grounded language fundamental zum Verständnis von Robotik als lohnt es sich die Thematik näher zu untersuchen.

Von der selbstbeschreibung her ist Grounded language eine Mischung als Sprachwissenschaft mit Informatik. Natürliche Sprache wird verwendet um ein Informatik-Problem z.B. Robotik-Steuerung zu lösen. Technisch gesehen ist das ein vielversprechender Ansatz allerdings ist unklar wo Literatur über diese Thematik einsortiert werden muss. Weder in die Linguistik noch in die Informatik passt grounded language wirklich hinein. Ein möglichers Gebiet wäre die Nachrichtentechnik welche sich mit der Informationsübertragung vom Sender zum Empfänger beschäftigt, nur leider besteht Nachrichtentechnik eher aus der technischen Realisierbarkeit also wie Bits über einen Kanal fließen und weniger in der semantischen Analyse einer Nachricht.

Vermutlich werden die meisten Informatiker noch nie etwas von grounded language gehört haben. Der Grund ist dass Informatik seine Wurzeln in den exakten Naturwissenschaften hat also verwand ist mit der Mathematik und der Physik. Um Computer zu bauen benötigt man Elektrotechnik und dort speziell  Transistoren. Um Computer zu programmieren benötigt man Algorithmen welche in der Mathematik untersucht werden. Leider hat grounded language mit beidem nichts zu tun. Es ist keine matghematik sondern es ist verwand mit der Sprachwissenschaft von Ferdinand de Saussure der untersucht hat wie Zeichen ihre Bedeutung erhalten. Die Methoden innerhalb der Sprachwissenschaft unterscheiden sich grundsätzlich von den Methoden in der Mathematik. Sprachwissenschaft wird als Geisteswissenschaft bezeichnet und gehört wie Geschichte und Soziologie zur Kulturhistorie.

Wollte man grounded language angemessen berücksichtigen müsste man eigentlich eine neue Kategorie erstellen zusätzlich zur bekannten Dewey Klassifikation. Das ist praktisch nicht durchführbar weil ja die idee hinter der etablierten Systematik darin besteht die Literatur auf dieses Raster einzurodnen. Um das Problem der Einordnung zu lösen muss man zuerst einmal grob definieren ob grounded language im Bereich Naturwissenschaft oder im Bereich Geisteswissenschaft veroret ist. 

Am ehesten könnte man grounded language als Naturwissenschaft bezeichnen und zwar weil der technisch-mathematische Aspekt im Zentrum steht. Es geht weniger darum Sprache an sich zu beschreiben sondern Sprache wird verwendet um Roboter zu steuern. Ähnlich wie Motion capture ist es damit im Bereich Naturwissenschaft -> Informatik -> Robotik lokalisiert. Auch bei Motion capture verfahren werden bekanntlich Ideen aus der Sportwissenschaft verwendet, allerdings ist Mocap zunächst einmal ein technisches Verfahren und wird daher von der Informatik definiert.

Bei grounded language ging es technikhistorisch immer darum, Sensordaten mittels Computer in textuelle ausgabe zu übersetzen. Frühe Beispiele waren das SHRDLU Projekt (Terry Winograd, 1968) oder Commentator scene description (Bengt Sigurd, 1980). Der Computer wurde also zwingend in diesen Projekten eingesetzt. Da das SHRDLU Projekt primär ein Artefakt der Informatik war, ist auch grounded language ein Teilbereich der Informatik-Geschichte.

Die klassische Sprachwissenschaft untersucht ebenfalls Sprache allerdings geht es um Sprache wie sie von Menschen oder Tieren verwendet wird, nicht um Sprache die von Algorithmen erzeugt wird. Sobald der Computer im Zentrum steht wird ein Thema als Informatik betrachtet. Im Fall von grounded language steht der computer zweifelsfrei im Zentrum der Betrachung. Es geht darum Sprache soweit zu formalisieren dass sie von Computern zur Interaktion verwendet werden kann. Computer sind definitionsgemäß innerhalb der Informatik beheimatet und werden nach naturwissenschaftlichen Prinipien beschrieben.

Mag sein dass für die Beschreibung von grounded language auf Theorien der Sprachwissenschaft und Psychologie zurückgegriffen wird, aber das könnte man über Computerspiele auch sagen. So verwendet das Spiel "Sim City" elemente der Architekturplanung während Malprogramme einen Starken bezug haben zur Kunst. Trotzdem sind diese Beispiel innerhalb der Informatik verortet weil der Computer jedesmal im Zentrum steht.

Grounded language kann man daher als neues Aufgabengebiet für einen Computer definieren. Anstatt nur Daten über ein Leitung zu übetragen wie das durch das Internet erfolgt und anstatt einfach nur Zahlen aufzuaddieren wie das mit einer Tabellenkalulation möglich ist, wird durch grounded language der Computer in die Lage versetzt natürliche Sprache an Roboter zu senden und zu empfangen. Und weil dies über Algorithmen funktioniert ist die Informatik die richtige Anlaufstelle für eine weitere Literaturrecherche.

Picture dictionary for kitchen domain

 

There are objects, activities and adjectives. Such a dictionary provides a vocabulary for talking about the subject. It allows a robot to localize the objects with a camera and understand basic requests like "wash dirty plate", "stir hot bowl", "bake cake in oven".

Some of the entries in the picture are labeled wrong, this is a technical problem. 

September 15, 2026

Das Symbol grounding Problem als Nischendisziplin

Die meisten Informatiker werden noch nie vom Symbol grounding problem gehört haben. Der Grund ist dass es sich nur schlecht in bisherige Wissenschaftsdisziplinen wie Mathematik, Elektrotechnik oder Informatik einordnen lässt. Gleichzeitig ist grounded language fundamental für die Steuerung von Robotern, so dass es Sinn macht die Thematik näher zu erläutern.

Im Kern des Symbol Grounding problem steht natürliche Sprache also Deutsch oder Englisch welche ni einem Wörterbuch gespeichert ist. Am ehesten gehört es also in die Linguistik welche Sprachen erforscht und dessen Bezug zur Wirklichkeit. Wörter referenzieren auf Sensormuster die ein Roboter detektiert sowie auf Kommandos die von einem Menschen formuliert werden. Die konkrete Interkation wird als Sprachspiel bezeichnet. Ein typisches Sprachspiele ist das "pointing game". Dabei sagt Person A einen Begriff wie "Raum B" und Person B muss auf diesen Ort zeigen.

Worte werden in einer numerischen Darstellung von Computern verarbeitet was als embedding bezueichnet wird. Damit sind die wichtigsten Elemente des Symbol grounding bereits erläutert. Weitere Details beziehen sich auf dieses Grundgerüst bestehend aus:
- Wörterbuch, Sprachspiel, embedding

Es verwundert wenig dass grounded language von der etablierten Informatik ignoriert wird, weil die Grundannahme lautet die Realität nicht über Zahlen sondern mittels Worten zu beschreiben. Für Worte gibt es keine mathematische Theorie sondern Worte werden nur außerhalb der Mathematik behandelt.

Worte und Sprachspiele dienen der Kommunikation also dem Nachrichtenaustausch. Man könnte also Grounded language als Teil der Nachrichtentechnik behandeln, wenn man es denn in etablierten Wissenschaftsdisziplinen erläutern möchte. Es gibt einen Sender, eine Botschaft und einen Empfänger. Damit ist zugleich der wesentliche Unterschied zu einem Algorithmus benannt. Ein Algorithmus wird auf einer Turing Maschine also einer CPU ausgeführt, während es in der Nachrichtentechnik keine Algorithmen gibt sondern es gibt Informationsübertragung, also einen Kanal auf dem Bits fließen.

Die klassicshe Nachrichtentechnik inkl. computer basierter Kommunikation ist ein gut erforschtes Gebiet. Deren größter praktischer Erfolg ist das Internet also ein Rechnerverbund der mittels TCP/IP Protokoll interagiert. Man kann sich grounded language als eine Art semantischer Nachrichtentechnik vorstellen wo neben dem Übertragen von Daten auch die Bedeutung der Daten von Bedeutung ist. Das zentrale technische Element ist ein parser, also ein Computerprogram was natürliche Sprache in einen numerischen Vektor übersetzt. Die Übersetzung von Sprache in Vektoren und von vektoren zurück in Sprache dient der Kompression. Ein sehr großer Zustandsraum wie er in der Robotik üblich ist wird projekziert auf einem kompakten selbst definierten Zustandsraum der vollständig durchsuchbar ist und von Computern verarbeitet werden kann. Damit lassen sich np harte probleme lösen.

Dazu ein praktisches Beispiel. Ein Roboter befindet sich in auf einer großen 2d Karte die aus 800x600 Pixeln besteht. MAthematisch gesehen kann der Roboter also einen der möglichen 480000 Pixel als Position einnehmen. Von dort aus kann er weitere Positionen erreichen wodurch sich die Zahl möglicher Trajektorien exponentiell erhöht. Aus Sicht von grounded language kann der selbe Roboter nur eine Position haben in [Raum A], [Raum B] oder [Raum c]. Es gibt also nur 3 mögliche Zustände des Systems. Jedes Label wie [Raum A) referenziert auf eine große Zahl von pixel in der realen Karte und fasst diese unter einem einzigen Begriff zusammen. Selbst wenn der Roboter über mehrere Hd Kameras plus Lidar Sensor verfügt, welche einen datenstrom im Gigabit Bereich liefert, bleibt sprachlich gesenen das System sehr überschaubar. 

August 27, 2026

Weiteres Beispiel zur Datenkompression mit grounded language

 Im vorherigen Blogpost wurde bereits ein Lagerroboter als Beispiel erwähnt. In diesem Post soll die Thematik Datenkompression vertieft werden.

Ausgangspunkt ist das Problem in der klassischen KI Forschung bis ca. 2010, dass ein Roboter in einem sehr großen Zustandsraum agiert der sich nicht effizient mittels vorhandener Hardware durchsuchen lässt. Diese Problemklasse wird als np hard problem bezeichnet und betrifft motion planning, senor perception und STeuerung von Robotern ganz allgemein.

Obwohl die Informatik über hunderte von Algorithmen, Programmiersprachen, und schnellen CPU ist verfügt ist keine Technologie mächtig genug Probleme mit einem sehr großen Zusttandsraum zu lösen. Damit ist Künstliche Intelligenz unlösbar.

Die einzige Ausnahme besteht darin, das Ausgangsproblem in ein niedrig-dimensionales Prolbem zu überführen, natürliche Sprache dient dazu als kompressionstechnik. Für den erwähnten warehouse roboter bietet sich eine Minivokablur an, was Zielorte und Ereignisse beinhaltet:

Zielorter: regalA, regalB, Korridor
Ereignisse: Kollison, Batterie_leer, ziel_erreicht

Die Vokabelliste definiert einen neuen Zustandsraum auf einem symbolischen Level. Er besteht aus 6 möglichen Worten und abstrahiert vom ursprünglichen Zustandsraum. Die Frage ist nicht länger wie man die Kamerabilder des Roboters speichert, oder wieviele Anzahl möglicher Trajektorien es gibt, sondern die Frage ist welche der 6 Wörter gerade aktiv ist.

Der neue sprachliche Zustandsraum kann viel leichter auf einem Roboter gespeichert werden. Man speichert die Vokabelliste in einer Tabelle und kann darauf referenzieren. Damit wird die ursprüngliche Problembeschreibung modifiziert. Es geht nicht länger darum einen Lagerroboter zu steuern der über eine hochauflösende Kameras und mehrere Servo-M;otoren verfügt sondern das neue problem ist, die ist situation des Roboter mittels Natürlicher Sprache zu beschreiben.

Eine Analogie aus der Nicht informatik ist eine Landkarte. Karten werden dazu verwendet größere Gebiete übersichtlich darzustellen. Ein Gebiet wie z.B. ein Wald sind auf einer Karte mit einem einfachen Symbol markiert. Obwohl der Wald über hunderte von Bäumen enthält sind diese nicht eingezeichnet sondern es gibt nur ein grünes Rechteck mit dem Symbol "Wald". Erst der Rezipient der Landkarte dekodiert die Information und schließt aus dem Symbol auf die ursprüngliche Realität. Dadurch reduzieren Landkartieren die Komplexität der Wirklichkeit.

Technisch gesehen lässt sich grounded language für Roboter erstuanlich einfach implementieren. Man extrahiert aus einer Szene zuerst Features und konvertiert diese dann in einen Text. Der Programmieraufwand in lines of code ist überschaubar und die benötigte CPU Leistung ist gering. Dennoch waren solche Systeme vor 2010 selten bis gar nicht vorhanden. Weniger aus technischen Gründen als vielmehr aus einem mangelnden Verständnis für das Symbol grounding problem. Bis 2010 war zwar bekannt, dass KI Probleme np hard sind und der state space zu groß ist um diesen zu durchsuchen, es war allerdings unklar, dass natürliche Sprache darauf die Antwort ist. Was stattdessen untersucht wurde, waren heuristiken, Reward Funktionen und sampling basierte Algorithmen wie RRT.

Mit diesen Verfahren konnte man Fortschritte bei motion planning realisieren, allerdings waren das unbedeutende Detailverbesserungen. Der Durchbruch erfolgte erst, durch Verwendung natürlicher Sprache als Abstraktionsmechanismus. 

Datenkompression mit grounded language an einem Beispiel

 Das Hauptproblem in der KI Forschung bis 1990 war das state space problem, also die Hohe Anzahl möglicher Zustände eines Systems. Das state space problem verhinderte das KI Probleme wie Motion planning von einem Computer in echtzeit gelöst werden konnten. Die vorhandenne Algorithmen waren nicht effizient genug und die vorhandene Hardware war zu langsam.

Die Antwort besteht in der Datenkompression mit Hilfe von grounded lanugage. Sprachw wird verwendet als Karte die über die Domäne gelegt wird. Dazu ein Beispiel: Angenommen die lagerhaus wird in einer 800x600=480000 pixel großen Übersicht gespeichert. In diesem Beispiel gibt es eindeutig ein state space problem weil die Anzahl von rund 0.5 Mio unterschiedliche Pixel die aus verschiedenen Farben bestehen eine sehr große Last erzeugt. Um diesen Rohdaten Objekte oder Wege zu erkennen bräuchte man Supercomputer. Mittels semantischer Datenkompression lässt sich die Aufgabe vereinfachen. Zuerst definiert man eine Vokabelliste (RegalA, RegalB, Korridor), dann definiert man Bereiche in der Karte auf die diese Vokabeln zutreffen. Man erhält dadurch eine annotierte 2d Karte.

In dieser neuen Realität ist das state space problem gelöst. Der Roboter kennt lediglich drei Begriffe "RegelA, RegalB, Korridor" und kann ermitteln wo er sich befindet. Durch eine Karte wurde also die hochkomplexe Wirklichkeit stark komprimiert und lässt sich leichter maschinenlesbar speichern.

Man kann also sagen, dass natürliche Sprache zur Datenkompression verwendet wird. Durch cid Vergabe von Begriffen werden 2d-Bereiche oder Events in der Ausgangsdomäne gelabelt. Diese Label dienen als Platzhalter wodurch Komplexität gesenkt wird. Der Boboter benötigt nicht länger die Information über die 800x600 Pixelkarte selber sondern er referenziert mit hilfe der Vokabelliste viel effizineter auf die wirklichkeit.

August 26, 2026

KI durch Kompression

Bereits 1973 hat James Lighthill erkannt, dass Künstliche Intelligence an der kombinatorischen Explosion scheitert. Gemeint ist, dass dass z.B. ein Lagerroboter einen sehr großen Zustandsraum besitzt mit Millionen von unterschiedlichen Aktionsmöglichkeiten. Diesen Zustandsraum mittels Computer zu durchsuchen ist mathematisch unmöglich, in der Informatik wird das als NP harte Problemklasse bezeichnet. Das Grundproblem, womit sich Generationen von KI Forschern konfrontiert sahen, war also die Suche in einem riesigen Zustandsraum. 

Die Antwort auf die Fragestellung besteht darin, natürliche Sprache als Kompression zu nutzen. Damit lässt sich der Zustandsraum eines Roboters verkleinern. Dieses Konzept ist als Symbol grounding problem bekannt und meint, dass der originale Zustandsraum bestehend aus Sensorwerten und Servomotoren-Signalen mittels natürlicher Sprache kartiert wird und dann von Computern verarbeitet wird.

Ein Lagerroboter hat nicht länger Millionen möglicher Trajektorien, die es durchzuprobieren gilt, sondern die Welt des Lagerroboters besteht aus einer Vokabelliste von weniger als 50 Worten womit er die Umgebung analysiert und Handlungen ausführt. Dieser diskerete Symbolvorrat reduziert den Zustandsraum und eine maschinelle Speicherung inkl. dem Planen von Handlungen wird möglich.

August 21, 2026

Piano movers problem with head up display including inner speech

  

The piano movers problem is one of the milestone subjects in the history of motion planning and was discussed frequently. Common knowledge until 2010 was that its an example for np hard problems, that the state space is very large and its very difficult to solve the problem with existing algorithms like RRT.

Instead of discussing the issue only from a mathematical and algorithmic standpoint there is need to introduce natural language as communication code between the low level task and a high level external oracle who gives advice what to do next. Such an interface can be realized as a head up display. The graphical area on top shows the classical piano movers problem as a 2d rendering, while the textual widget on the bottom shows the comments from the external oracle which monitors the scene.

The external oracle can be located outside of the robot or it can be embedded inside the robot than its called the inner voice of the robot. In the example it generates speech for observation, goal and action. This translates the piano movers problem into a an abstract textual problem similar to an interactive fiction story. 

August 20, 2026

Wie man grounded language implementiert, Schritt für Schritt

Zuerst definiert man eine Vokabellliste die für das Problem angemessen ist. Bei einem Minimalbeispiel soll ein Roboter sich auf einem Graph bewegen. Typische Vokabeln sind:
nord, osten, süden, westen, rot, grün, blau, wegpunkt, schnell, langsam, stop

Im nächsten Schritt wird ein Bild zu Text Parser programmiert, dieser erzeugt eine textuelle Beschreibung für die aktuelle Situation. Eine Beispielausgabe wäre:
"Roboter ist im Norden, nahe des grünen Wegpunktes. Er bewegt sich schnell"
oder "Roboter ist im Süden, nahe des roten Wegpunktes. Er hat gestoppt".

Selbstverständlich ist diese Textausgabe nicht sehr eloquent und basiert auf einer einfachen Template die mit aktuellen Werten gefüllt wird. Aber das ist akzeptabel, wichtig ist dass überhaupt eine Textausgabe stattfindet.

Um den umgekehrten Kommunikationsweg von einem Kommando zu einer Aktion zu implementieren benötigt man eine Reward Funktion. Ein Kommando wie "gehe nach NOrden" wird in einen SCore umgerechnet. Wenn der Roboter im Norden angekommen ist beträgt sein Score 1.0 wenn er weit davon entfernt ist nur 0.0".

Damit stehen alle wichtigen Elemente zur Verfügung um mit dem Roboer in natürlicher Sprache zu kommunizieren. Der Roboter beschreibt die aktuelle als Text und reagiert auf Kommanddos. Zumindest für die eingeschränkte Domäne des "Roboter auf einem Graph" Problems ist damit eine Kommunikation zwischen Roboter und menschlichem Bediener möglich.

Mit Hilfe dieser Kommunikation kann man komplexe Aufgaben automatisieren. man sendet z.B. eine Befehlsfolge an den Roboter:
"1. fahre nach Norden, suche den grünen WEgpunkt, warte dort.
2. dann fahre nach Süden zum Roten Wegpunkt. und zwar schnell"

Diese Kommandos übersetzt der Roboter zuerst in einen Reward und nutzt den Reward dann um Aktionen auszuführen. Während er das tut beschreibt der Roboter in der Logdatei was er gerade macht also wo er ist, was er sieht und was die nächsten Schritte sind.



August 18, 2026

History of the symbol grounding problem

 Symbol grounding is a multidisciplinary approach which has evolved over years. The follwoing timeline shows the milestones during the development. First important innovation was the invention of written language thousands of years ago, in the year 1844 the morse code was invented to transmit signs over distance. In 1968, the SHDRLU project was initiated which was a text to graphics interface on a computer. In 1991 Rodney brooks recognized that robots need to interact with the environment, in 2003 a voice controlled robot was built by Deb Roy and in the year 2023 the Lingo-1 Vision language model was published for controlling self driving cars.

The shared similarity between these innovations is the focus on natural language for connecting two systems. A speaker sends a command to a hearer. Implementing such an interface on a computer allows to teleoperate a robot with language. This is the precondition for automation. 

The crucial building block for symbol grounding is natural language which is used as a reference system to the external world. Thanks to predefined words for objects and activities its possible to label sensory data and receive high level commands. A robot can recognizhe that he stands in front of an obstacle, and a human operator can submit a commoand to the robot like "move around the obstacle to the left". It took decades and even centuries until human researchers have discovered the power of natural language and were able to implement textual interface for robots. Its likely that future robots invented in 20 years from now will work with the same principle.

3300 BC,Cuneiform writing system in Mesopotamia
1500 BC,sundial showing the time of the day
600 BC,Latin alphabet available in Italy
322 BC,correspondence theory of truth by Aristotle
1386,Salisbury Cathedral tower clock with a bell
1440,printing press by Johannes Gutenberg
1505,Pomander Watch by Peter Henlein
1792,optical telegraph by Claude Chappe
1844,morse code by Samuel Morse
1870,Engine Order Telegraph by William Chadburn
1876,commercial typewriter by Remington
1878,chronophotography "The Horse in Motion" by Eadweard Muybridge
1903,Telekino remote controlled boat by Leonardo Quevedo
1915,Therblig notation by Frank Gilbreth
1915,rotoscoping animation technique by Max Fleischer 
1920,AAC Communication board by F. Hall Roe
1928,Labanotation dance notation by Rudolf von Laban
1929,Televox robot by Westinghouse
1930,motion tracking by Nikolai Bernstein
1949,Turing test by Alan Turing
1954,Georgetown-IBM Experiment with russian translation
1959,Pandemonium architecture by Oliver Selfridge
1962,ANIMAC motion capture by Lee Harrison III
1963,ASCII code
1966,ELIZA chatbot by Joseph Weizenbaum
1968,SHRDLU natural language understanding by Terry Winograd
1971,Lexigram for communicating with apes by Ernst von Glasersfeld
1977,Zork I text adventure by Tim Anderson
1977,Tour model instruction following by Benjamin Kuipers MIT AI lab
1980,Chinese room argument by John Searle
1980,Commentator scene description by Bengt Sigurd
1980,Finite State machine in Pacman videogame by Tōru Iwatani
1981,Karel the robot programming language by Richard Pattis
1983,MIDI music protocol
1983,M.I.T. Graphical Marionette by Delle Maxwell
1984,Castle Adventure by Kevin Bales
1987,Maniac Mansion point&click adventure by Ron Gilbert
1987,Vitra visual translator by Wolfgang Wahlster
1989,Speech activated manipulator SAM by Michael Brown
1990,Physical Grounding Hypothesis by Rodney Brooks
1990,paper "The symbol grounding problem" by Stevan Harnad
1991,paper "Intelligence without Representation" by Rodney Brooks
1993,AnimNL computeranimation by Norman Badler
1993,conceptual spaces by Peter Gardenfors
1994,Abigail scene recognition by Jeffrey Siskind
1996,Interaction machines by Peter Wegner
1998,Rocco Robocup commentator by Dirk Voelz
1999,trec-8 Text Retrieval Conference
2000 Kismet social robot by Cynthia Breazeal M.I.T.
2003,M.I.T. Ripley robot by Deb Roy
2006,Marco route instruction following by Matt MacMahon
2007,Simbicon computer animation by Michiel Panne
2010,Motion grammar by Mike Stilman
2010,M.I.T. forklift by Stefanie Tellex
2011,IBM Watson Question answering by David Ferrucci
2013,Word2vec algorithm by Tomas Mikolov
2015,Poeticon++ trajectory recognition by Yiannis Aloimonos
2015,DAQUAR VQA dataset by Mateusz Malinowski
2017,Transformer Architecture by Ashish Vaswani
2020,Vision language model by different authors
2023,Wayve Lingo-1 self driving car

August 08, 2026

Sprachspiele im Kontext von Robotik

Die Erforschung künstlicher Intelligenz drehte sich lange Zeit um die Frage welche Art von Softwareframework, Algorithmus oder Programmiersprache benötigt wird um denkende Maschinen zu realisieren. Die Annahme hinter dem Cyc Projekt von Douglas Lenat lautete, dass eine Ontologie der Grundbaustein sei, das Cam-Brain Projekt von Hugo de Garis unterstellte dass im Kern ein neuronales Netz benötigt wird während Edward Feigenbaum vermutete dass sich Künstliche Intelligenz mit Hilfe von Expertensystemen realisieren läst.

Im Laufe der Zeit wurden sehr viele gegensätzliche Technologien und Annahmen entwickelt um Künstliche Intelligenz zu verwirklichen. Ein Sprachspiel im Sinne von Ludwig Wittgenstein ist nur ein weiterer Vorschlag unter vielen. Dennoch lohnt es sich, das Thema nähter zu untersuchen. Weil das Prinzip von Sprachspielen deutlich anders ist als z.b. ein Expertensystem oder ein neuronales Netz.

Sprachspiele sind ein formalisiertes Interface zwischen interner und externer Realität von Systemen. Es ist ein Test, ob die Spielteilnehmer Sprache erzeugen und verstehen mit derer man die externe Realität beschreibt. Ein typisches Beispiel ist das "Ich sehe was du nicht siehst Spiel". Dabei beschreibt Person A ein Objekt aus der Realität anhand von Eigenschaften, er formuliert aussagen wie "das Objekt ist rund, steht in der Küche, hat Beine" und daraus folgert Person B "es ist ein Tisch".

Die Gemeinsamkeit von allen Sprachspielen ist, dass Sprache in einem interkationsspiel genutzt wird um auf die Wirklichkeit zu referenzieren. Es geht immer darum, Dinge abzufragen die in der Realität vorkommen, Anweisungen geben was in der Realität zu tun ist oder sonstwie Sprache und Wirklichkeit in Beziehung zu setzen. Sprachspiele sind eine gute Möglichkeit eine Fremdsprache zu erlernen und dienen dazu den Wissensstand zu erfassen. Wenn eine Person oder ein Computer in einem Sprachspiel eine hohe Punktzahl erreicht, ist diese Person mit der jeweiligen Sprache vertraut.

Anders als die eingangs erwähnten Technologie zur Realisierung von künstlicher Intelligenz wie Onotologien oder Expertensysteme sind Sprachspiele nicht technisch definiert sondern haben ihren Ursprung in der Philosophie und der Linguistik. Es geht um Themen wie interaktivität, Sprache, Wirklichkeit. Die Annahme lautet dass innerhalb dieser Begriffe Künstliche Intelligenz möglich ist, das also denkende Maschinen ein Interface zwischen internem System und externer Realität sind.

Seit den 1980er wurden Brettspiele als Testumgebung für denkende Maschinen genutzt. Schach ist ein sehr altes Beispiel für das Künstliche Intelligenz realisiert wurde, aber auch andere Spiele wie Tic Tac Toe, Dame, Backgammon und Go sind klassische Umgebung zur Erforschung von KI Algorithmen. Leider haben diese Brettspiele den Nachteil dass sie nicht gut nach oben skalieren. Ein Computer der perfekt Schach spielt ist nicht automatisch im Stande einen Roboter zu steuern. Deshalb eignen sich die erwähnten Brettspiele nur sehr eingeschränkt dazu KI näher zu erforschen.

Sprachspiele kann man als neuartiges Gesellschaftsspiel verstehen was ähnlich wie Schach Regeln folgt aber viel besser nach oben skaliert. Ein Computer der das "Guess what" Sprachspiel beherscht ist zugleich auch in der Lage mit diesem Wissen einen Roboter zu steuern. Scheinbar können Sprachspiele die Kernidee von Künstlicher Intelligenz viel besser formalisieren als frühere Brettspiele.

July 25, 2026

Engine for grounded language

One possible explanation why the symbol grounding problem has emerged late in the history of computer science is because the theory is difficult to realize on a computer. Suppose natural language is important for robot control, the problem is create a language parser which is working for a concrete domain.

A possible command for a robot might be "Move until obstacle and then stop". Each of the words is stored as a string, but it remains unclear who to process the instruction into actions for a robot. The reason is, that the sentence is formulated in English but computers need a programming language as input. Even if every word is encoded as a number, it doesn't make sense to submit an array with numbers to the robot because its not possible to add or subtract the values in a meaningful way.

In general the problem is how to convert natural language into a computer program. Without solving this issue, the symbol grounding problem remains only a philosophical problem without any practical consequences.

The good news is, that the problem of programming a parser can be solved. Not with tools from computer science but by using techniques from linguistics, namely language games. Instead of treating language parsing as an algorithm problem, the idea is to invent around words a puzzle game. Typical language games are:

- Name guessing game. Player1 points to an object in the reality, and Player2 has to tell the name
- NPC quest game, a non player character in a role playing game formulates a quest like "bring me the sword from the wood" and the player has to fulfill the task
- bounding box game, player1 says a word like "table" and player2 has to draw a bounding box around this object

All these language games are located outside of computer science. They have nothing to do with algorithms, programming language nor existing robotics libraries, but they are games played with 2 human players.

The interesting situation is, that its possible to simulate the games with a computer. The software encodes the rules of the language game, determines the score for the human player and then the player can take action inside this game.

The problem is not how to program a certain parser, but the problem is how to formulate the game outside of a computer first. A well formulated game can be implemented in a software with ease. The programmer needs only the specification of the game including its rule, and then its possible to program the game with python. The only requirement is, that the computer works like the original language game. The programming workflow is identical to implement card games and board games on a computer.



July 24, 2026

Practical demonstration of the Total turing test

Stevan Harnad coined around the year 1990 the term "total turing test" which is a philosophical description of an instruction following task in robotics. What is missing is a practical demonstration of such a test for a real robot.

Such a demonstration can be realized in a simple video game modeled as a language game. An entry level example is a navigation task in a graph. There are 8 nodes connected with lines and the robot has to move along the graph to reach a certain goal node. Possible interaction with the robot would be:
- what is your position?
- what is your battery status?
- which nodes are reachable from current position?
- Move north
- move to node #3.
- what is the shortest path to reach node #6?

The robot is in charge to answer these requests in natural language and with motor actions. The problem is easy enough to get implemented as normal computer code without using advanced large language models or vision language action models. The human to robot interaction can be simplified by using a codebook. The amount of possible commands is given in the menu and the human can select one of the commands. IN other words, the Total Turing Test (TTT) is some sort of speaker to hearer interaction game played between a human and a robot.

July 21, 2026

Head up display for a kitchen robot

The picture shows an artist version of a head up display. It contains of:
- camera picture of a kitchen
- text box with inner voice
- bounding boxes
- labels for the bounding boxes

Surprisingly, the information in the picture can solve the symbol grounding problem because the head up display connects visual perception with textual information. The text from the inner voice like "I need to find 200g of flour" can be converted into meaning with the help of the bounding boxes. There is a box available with such an ingredient. The task for the robot is not to plan actions but the main problem is to connect language from the inner voice with detected objects in the camera.

Such a link of visual objects with textual labels is the core element in grounded language. If the robot is able to identify objects from the text box, its possible to generate all sort of inner voice. For example, the robot can say that he needs to peel the banana or "open the oven". All these nouns and verbs are translated into position of the bounding box on the screen which allows to execute the action physically.

June 02, 2026

Grounding mechanism 1o1

 A DIKW pyramid consists of abstraction layers like Data, information and other. A grounding mechamism maps the items in the layer. In an example warehouse robot, the data layer cosnsits of sensor readings like GPS Coordinates, lidar distance, and battery capacity while the information layer consists of [tags] like "battery_full, north, obstacle_ahead".

The grounding mechanism generates the links between the entries. For example the lidar distcance of 10 cm is mapped to "obstacle_ahead" while the battery level of 10% is mapped to "Battery_empty".

In general, a grounding mechanism is some sort of matching game. it answers the question which situation is mapped to which description. Such a mapping is the core element of an advanced artificail intelligence.

To demonstrate why a matching game enables artificial intelligence let us assume an example. Suppose the human operator submits a command to the warehouse robot which is "move to the green area, grasp the small box on the left side, bring the box to the blue area, drop it into the shelf, then recharge your battery".

If the grounding mechanism is missing or was deactated, the command is interpreted as string with 144 characters. It wasn't formulated in the C/C++ programming langauge but it can be stored only in the main memory.

Suppose the robot has a builtin grounding mechanism, than its possible to parse the sentence word by word. The word "green" is matching to a certain RGB value, the word "box" is mapped to a certain shape in the camera, the word "shelf" is mapped to a picture of the shelf and so on. The parsing algorithm fetches a word from the sentences, and takes a lookup into the database to identify the item from the data layer of the DIKW pyramid. Understanding a sentence from a robots perspective has to do with matching items from the information layer to the data layer.

June 01, 2026

Symbol grounding problem as answer to np hard algorithms

 Before its possible to describe grounded language there is a need to explain who artificial intelligence was imagined until the year 1990. It was treated similar to computer programming in the sense that there is a CPU which executes a program and its up to the programmer to make the algorithm as intelligent as possible. Artificial intelligence was thought as a very advanced computer programmed which is executed by a computer.

In other terms, the computer was seen as a problem solving machine and the only detail problem was which sort of algorithm is needed to solve a certain problem. For example motion planning in robotics was solved with motion planning algorithms while computer chess was solved with alpha beta prunning algorithms. Most of these AI related algorithms were designed as search algorithms. The computer was used to traverse the state space of the domain and this allowed the computer to find the optimal action.

The symbol grounding problem formulated by Stevan Harnad questions this algorithm oriented paradigm. This might explain why even today grounded language is a niche topic within computer science. Because computer science and algorithms were often treated as the same thing, it was outside of the scope how to program a computer without an algorithm.

Let us listen closely how Harnad, Brooks and Steels are arguing about grounded language. The core element is the sensory perception of a robot. The assumption is that the perception is transmitted to the computer. There is no need to calculate something but the focus on the data transfer. A light sensor detects light and the information from the sensor is send over a cable to the computer. The symbol grouding problem doesn't focus on the computer itself, but on the cable between a sensor and a computer, very similar to a computer network. Computer networks are different from a turing machine, they are never running algorithms, but a computer network communicates data often organized in a protocol layer.

The paradigm shift from algorithm centric computers towards protocol oriented data transmission is the core element of the symbol grounding problem. Artificial Intelligence isn't explained as processing or program executation, but Artificical Intelligence is imaged as the air gap between two hosts.

Let us compare the hardware. In classical algorithm oriented AI the basic building block is a central processing unit, which can be a 32bit CPU. The CPU is built with transistors on a chip and gets controlled by Assembly language. In contrast, the symbol grounding problem assumes that there is a Cat5 copper cable which delivers packets. Its up to the network engineer to define the protocol of the packets.

The paradigm shift can be explained for np hard problems. NP hard is a certain category of problems related to artificial intelligence which can't be solved with a computer. Nearly all robotics motion planning problems like the piano movers problem or model predictive control are np hard. The term np hard is referencing to the runtime of an algorithm executed on a cpu. In other words, even a modern 64bit CPU can't solve these problems because the hardware is too slow.

The holy grail in computer science is how to solve np hard problems. The answer was given by Stevan Harnad in his famous 1990 paper. He didn't mentioned np hard problems, but its possible to solve np hard problem with grounded language. Instead of using a CPU to calculate a mathematical problem, a copper cable is used to solve a data transmission problem. This new perspective is powerful enought to solve motion planning problems in robotics.