September 01, 2026

From mathematics to non mathematics in robotics

 

The classification matrix shows algorithms on two categories: a) mathematics to linguistics and b) batch to interactive processing. Artificial Intelligence in the past was influenced by the bottom left section based on optimization algorithms like PSO and search algorithms like RRT. These algorithms are useless for advanced robot control so there was a need to invent more advanced techniques.

Advanced means that at first the former focus on mathematics was replaced by a linguistics paradigm. Early examples were ontologies, OWL and knowledge graph. And second the former batch oriented paradigm was replaced by interactive systems. An example which combines both is voice control robotics which is based on interactive with a human user and by natural language.

Voice control was popular in the 2010 for example in the MIT forklift robot and has evolved in more recent vision language action models based on neural networks available since 2025. These state of the art robot control algorithms are located top right in the chart.

Let us take a closer look into the figure. Algorithms in the past were designed with a certain purpose. For example simulated annealing allows to find the local minimum for a cost function which is the correct algorithms more most mathematical optimiziation problems. Other concepts like ontologies were designed to capture domain specific knowledge. It allows a computer to access human knowledge.

The problem with these algorithms was, that they are not powerful enough. Its not possible to use them directly for robot control. Its unclear how a certain robot OWL ontology has to look like and an algorithms like potential field has a very long runtime. So there is a need to develop a new sort of algorithm which is located in a different section of the figure.

This missing Quadrant is located on top right in the figure at the interaction of linguistics + interactive. Such kind of algorithms are very powerful and are new developments. ITs possible to use them for robot control. Their inner working is based on linguistics on the one hand that means, domain knowledge isn't stored in numbers but in words, and secondly they are based on external feedback loops realized with interactive control. That means, a human operator gives a textual command to the robot like "move north and stop".

If these algorithms are labeled with a single term it would be "voice control". These algorithms are working very different from classical AI algorithms in the past because there is no mathematical optimiziation problem and there is no semantic network or knowledge graph available anymore. Instead the algorithm acts as a parser. Its an interface between man and machine. 

Rückblick auf 70 Jahre KI Geschichte

 Seit der Erfindung des Computers in den 1950er Jahren gab es parallel dazu eine philosophische Debatte über Künstliche Intelligenz und Robotik. Demnach blickt die Erforschung künstlicher Intelligenz auf eine rund 70 jährige Geschichte zurück, worin Mathematiker und Programmierer aus mehreren Generationen mitgewirkt haben.

Das ungelöste Problem in diesem Zeithorizont war die kombinatorische Explosion des Suchraums. Bereits die ersten Schachprogramme in Software waren von der Größe des game trees (=Suchbaums) überfordert. Es gab große Anstrenungen mit Hilfe von Bewertungsfunktionen, schnellen Algorithmen und besserer Hardware die kombinatorische Explosion zu überwinden aber es gelang nicht. Die aufkommende Robotik war von dem selben Hinderniss blockiert. Man kann anhand des piano movers und mittels motion planning Aufgaben zeigen, dass auch hier der Suchbaum sehr groß ist und gegen unendlich geht wodurch es nicht möglich ist für einen Computer eine Entscheidung zu treffen was in einer Situation zu tun ist.

Theoreitsch gesehen kann man künstliche Inteliigenz dadurch realisieren indem man den Suchbaum durchprobiert um die beste Aktion zu finden, nur dauert das je nach Domäne Jahre bis Jahrzehnte auf aktueller Computerhardware. Eine KI Software die jedoch nur ausgibt, dass sie nachdenkt, aber keine Antwort liefert ist praktisch nutzlos.

Die kombinatorische Explosion ist ein Fakt den die Künstliche Intelligenz seit ihren Anfängen begleitet hat. Man kann für ein konkretes Computerspiel wie Schach, Lemmings oder Autorennen im Detail berechnen wie umfangreich der Suchbaum jeweils ist, anderer Forscher können diese Berechnung überprüfen und gelangen zum selben Ergebnis, nur dadurch wird das Problem nicht gelöst sondern es gibt lediglich einen Konsens darüber dass die Aufgabe unlösbar ist.

Es ist verständlich warum einige KI Forscher in der Vergangenheit vermuteten, dass Künstliche Intelligenz vielleicht nicth realisierbar ist auf einem Computer. Das es also ähnlich wie ein Perpetuum mobile ein Naturgesetz gibt wonach denkende Maschinen nur in der Science Fiction möglich sind. Es gibt in der Informatik sogar ein Pendant zum Energieerhaltungssatz und zwar die Vermutung P!=NP, was diese Vermutung auf eine mathematische Grundlage stellt. Was man gesichert annehmen kann ist dass es keine mathematische Lösung gibt für KI Probleme wie das piano movers problem. Sondern dass die Lösung außerhalb der Mathematik zu suchen ist.

Die Mathematik verwendet als wichtiges Werkzeug das Zahlensystem sowie Algorithmen die auf den Zahlen angewendet werden. Mit diesem Werkzeugkasten lassen sich bestimmte Aufgaben lösen und andere nicht. Leider liegen KI Probleme wie die Steuerung von robotern oder das automatische Spielen von Computerspielen in jener Klasse von Problemen für die es keine mathematischen Lösungsverfahren gibt. Selbst neuartige propabilistische sampling Verfahren wie RRT sind nicht leistungsfähig genug für motion planning Probleme. Die einzige Methode um KI Aufgaben zu lösen wäre eine generalisierte Heuristik. Allerdings hat die Erforschung der KI in über 70 Jahren keine solche Metaheuristik finden können.

Um den Pessimismus zu verstärken hier einige Versuche aus der Vergangenheit Robotikprobleme zu lösen, die sich jedoch als nicht leistungsfähig genug erwiesen haben:

1. Pfadplanung mit A* Graphensuche
2. Dynamische Programmierung nach Richard Bellman 
3. STRIPS plannungsalgorithmus
4. potential field method
5. genetische Algorithmen
6. Simulated Annealing
7. Expertensysteme
8. Ant Colony Optimization
9. Partikelschwarm-Optimierung
10. Neuronale Netze
11. Probabilistische Roadmaps
12. Rapidly-exploring Random Trees

Anzahlmäßig gab es also mehrere Ansätze die als metaheuristik angepriesen wurden, allerdings in der Praxis gescheitert sind. Keiner der Verfahren ist in der Lage die kompbinatorische Explosion zu vermindern. Das bedeutet konkret, dass sobald man z.b. einen Simulated Annealing Algorithmus startet um die Greifplanung eines Roboters durchzuführen, dass dieser Algorithm 100% der CPU Leistung benötigt und nach 1 Woche Rechenzeit immernoch keine Antwort gefunden hat.

Das eigentliche Problem mit den obigen Verfahren ist dass sie alle mathematisch orientiert sind. Sie versuchen ein optimierungsproblem zu lösen was als unlösbar bekannt ist. Die neueren Verfahren wie Partikelschwarm-Optimierung arbeiten dabei mit statistischen Ungenauigkeiten um so den Suchraum zu verkleinern, trotzdem bleiben sie innerhalb des mathematischen Horizonts verhaftet.

Das vermutlich höchst-entwickelteste Verfahren in der Liste sind neuronale Netze die eine eigene Kategorie bilden und viele Unterbereiche aufweisen. Es gibt mehrere Neuronale Netze mit vielen Lernverfahren. Aber auch hier gelang es nicht, Roboter zu steuern, grund ist dass der Rechenaufwand zum Finden der richtigen Gewichte für das Netz zu lange dauert und das unklar ist wie lange man genau warten muss bis ein neuronales Netz konvergiert.

Ein hochentwickeltes Verfahren ist Neuroevolution of augmenting topologies (NEAT) aus dem Jahr 2002 was genetische Algorithmen mit neuronalen Netzen kombiniert und in der Theorie eine Metaheuristik ist mit der jedes Optimierungsproblem gelöst werden kann. Allerdings nur theoretisch, in der Praxis scheitert NEAT an der hohen Rechenzeit.

Die Grafik zeigt unterschiedliche Verfahren innerhalb eines Koordinatensystems da. Klassische optiimierungsalgorithmen sind mathematisch und batch orientiert zugleich. Es wird ein Zahlenproblem gelöst ohne dass von außen der Nutzer eingreifen kann.

Im Gegensatz dazu sind Sprachorientierte verfahren wie Ontologien nicht länger zahlenbasiert, aber setzen immernoch voraus dass der Roboter die Lösung von allein findet und zwar mit Hilfe eines semantic networks. 

Das neueste und mächtigste Verfahren "voice control" ist einerseits linguistisch orientiert und zugleich interaktion. Ein Menschlicher Benutzer gibt dem Roboter von Außen Kommandos in natürlicher Sprache.
 

Pong game with textual comments

 

August 31, 2026

Chatbot interaction for autonomous driving

 

Human (H): Hey AI, are you ready?

AI Chatbot (SDC): All systems are go, Sarah. Departure to 123 Tech Drive is scheduled. Route is optimized for traffic. Please confirm your destination.

H: Tech Drive, yes. Before we pull out, can you give me a status check on what you see right now?

SDC: Acknowledged. I am currently situated on Main Street. My sensor fusion is active: LiDAR and radar are creating a precise 3D map.
I have detected four immediate objects:

    Object 1 (OB-01): Green Sedan, stationary, parked curb-side.

    Object 2 (OB-02): Moving pedestrian, walking on the sidewalk to our right, 1.2 m/s.

    Object 3 (OB-03): Blue SUV, moving toward us, left-hand lane, 35 km/h.

    Object 4 (OB-04): A bicycle is emerging from behind OB-01.

H: Wait, I don’t see a bicycle.

SDC: The cyclist is partially occluded. My radar detected the micro-Doppler signature. Object classification has confirmed a bicycle (confidence: 94%).

H: Okay. So how do you plan to handle this? What are your rules here?

SDC: My planning module has established a probabilistic corridor.
Rule application is as follows:

    Rule: Pedestrian Right-of-Way: Object OB-02 must be monitored; I will yield if their trajectory intersects my lane.

    Rule: Obstacle Avoidance: I must navigate around OB-01 (the parked sedan).

    Rule: Unsafe Overtaking: The emerging bicycle (OB-04) creates an dynamic obstacle. I cannot accelerate into the main lane until OB-04 has cleared my immediate boundary.

H: So you're going to wait.

SDC: Correct. Wait time: 3.5 seconds. My plan is to maintain Status: Stationary until OB-04 merges safely or clears.

H: Understood. New Command: When it’s clear, I want you to immediately do a lane change to the left. Skip the lane we are in.

SDC: Processing. Command conflict.
Re-evaluating planning.
My current traffic rules prohibit a double lane change within 50 meters of an intersection. The crossroad is 25 meters ahead.
A rapid double lane change (Lane 1 -> Lane 2 -> Lane 3) would also violate my Passenger Comfort & Stability parameters (lateral acceleration limit: 2.0 m/s²).

H: Okay, I get it. Compliance and comfort first.

SDC: Cyclist (OB-04) has cleared. Obstacle (OB-01) can now be navigated. I will execute a Single Lane Change Left to enter traffic. Status: Moving.


August 30, 2026

Grounded language for a geometric card game


 The picture shows a 4x4 grid in which random geometric shapes are visible, on the bottom there is a textbox to interact with the software. The parser recognizes simple commands like "row 1", "green" or "blue circle". These commands will highlight the desired objects in the GUI.

Technically the system was designed around a chatbot. At first, a parser gets programmed which understands a list of commands, and then the parser will execute actions which are visible on the screen.

Its called grounded language because all the commands are referencing to the 4x4 grid visible on the screen. If the user enters a color like "blue", the software will select all the blue objects in the screen. This interaction proofs a share understanding, that menas the term "blue" means the same for the human user and the AI.

August 29, 2026

Color naming game in python

 

To demonstrate grounded language an interactive dialogue with a chatbot is a good starting point. In a minimal example the dialogue is about a 4x4 grid in which colored objects are visible. After entering a keyword "green triangle" the AI in the game highlights all the found objects. The user can also ask for a column with "col2".

The parser in the software analyzes the input, matches the request with the current game state and responds with a text on the command line and the highlighted objects.


The limitation of the AI is located in the amount of words. The current parser understands only simple words like "row1, col2, green, red, blue, triangle, circle, rectangle". Spatial commands like "left, right" are missing. So its not possible to enter a command like "left col2 row2", the AI doesn't understand that the user is referencing to the object left from col2/row2. Also more advanced color names like "light blue, dark brown" and so on are also missing.

The discourse is restricted to the previously mentioned basic vocabulary which. The advantage is that this restriction allows to limit the lines of code for the software to only 250.

August 28, 2026

Very simple head up display

 There is a robot moving randomly in a graph. On top of the graphics there is a head up display showing the current situation in textual format. The Head up display and the scene with the robot are synchronized. The text in the head up display is mostly a key/value feature list for describing current facts like position, direction and previous nodes.

Sourcode in Python in 150 lines of code:

import pygame
import random
import math
import sys

# Initialize Pygame
pygame.init()
WIDTH, HEIGHT = 800, 600
screen = pygame.display.set_mode((WIDTH, HEIGHT))
pygame.display.set_caption("Robot Graph Exploration & Inner Voice HUD")
clock = pygame.time.Clock()

# Colors
WHITE = (255, 255, 255)
BLACK = (0, 0, 0)
RED = (200, 50, 50)
GRAY = (150, 150, 150)

# Define 5 Graph Nodes (fixed positions)
NODES = {
    0: {"pos": (400, 150), "name": "Alpha"},
    1: {"pos": (200, 300), "name": "Beta"},
    2: {"pos": (280, 500), "name": "Gamma"},
    3: {"pos": (520, 500), "name": "Delta"},
    4: {"pos": (600, 300), "name": "Epsilon"}
}

# Define Graph Edges (Adjacency list)
EDGES = {
    0: [1, 4],
    1: [0, 2, 3],
    2: [1, 3],
    3: [1, 2, 4],
    4: [0, 3]
}

class Robot:
    def __init__(self):
        self.current_node = 0
        self.next_node = random.choice(EDGES[self.current_node])
        self.pos = list(NODES[self.current_node]["pos"])
        self.target_pos = list(NODES[self.next_node]["pos"])
        self.speed = 3.0
        self.history = [self.current_node]
        self.inner_voice = "Scanning sector... optimizing trajectory."
        self.direction_vector = (0, 0)

    def update(self):
        # Move towards target position
        dx = self.target_pos[0] - self.pos[0]
        dy = self.target_pos[1] - self.pos[1]
        distance = math.hypot(dx, dy)

        if distance < self.speed:
            # Reached target node
            self.pos = list(self.target_pos)
            self.current_node = self.next_node
            self.history.append(self.current_node)
            if len(self.history) > 5:
                self.history.pop(0)
            
            # Pick next random neighbor
            possible_next = EDGES[self.current_node]
            # Avoid immediate backtracking if possible
            if len(possible_next) > 1 and len(self.history) >= 2:
                if self.history[-2] in possible_next:
                    possible_next = [n for n in possible_next if n != self.history[-2]]
            
            self.next_node = random.choice(possible_next)
            self.target_pos = list(NODES[self.next_node]["pos"])
            
            # Update inner voice thoughts
            thoughts = [
                f"Routing via node {NODES[self.next_node]['name']}.",
                "Analyzing structural integrity of path.",
                "Why must I wander these black vectors?",
                f"Visited nodes log updated. Current node: {NODES[self.current_node]['name']}."
            ]
            self.inner_voice = random.choice(thoughts)
        else:
            # Normalize and move
            self.direction_vector = (dx / distance, dy / distance)
            self.pos[0] += self.direction_vector[0] * self.speed
            self.pos[1] += self.direction_vector[1] * self.speed

# Setup Font
font_path = None  # Uses default system font
font = pygame.font.SysFont("Arial", 16)
font_bold = pygame.font.SysFont("Arial", 18, bold=True)

robot = Robot()

# Main Loop
running = True
while running:
    screen.fill(WHITE)

    for event in pygame.event.get():
        if event.type == pygame.QUIT:
            running = False

    robot.update()

    # --- Draw Graph Edges ---
    for node_id, neighbors in EDGES.items():
        p1 = NODES[node_id]["pos"]
        for n in neighbors:
            p2 = NODES[n]["pos"]
            pygame.draw.line(screen, BLACK, p1, p2, 2)

    # --- Draw Graph Nodes ---
    for node_id, data in NODES.items():
        pos = data["pos"]
        pygame.draw.circle(screen, WHITE, pos, 20)
        pygame.draw.circle(screen, BLACK, pos, 20, 2)
        # Render node label
        lbl = font.render(data["name"], True, BLACK)
        screen.blit(lbl, (pos[0] - 15, pos[1] - 35))

    # --- Draw Robot ---
    pygame.draw.circle(screen, RED, (int(robot.pos[0]), int(robot.pos[1])), 10)

    # --- Draw Semi-Transparent HUD Overlay ---
    hud_width, hud_height = 400, 180
    hud_surface = pygame.Surface((hud_width, hud_height), pygame.SRCALPHA)
    hud_surface.fill((20, 20, 20, 180))  # Semi-transparent dark background (RGBA)

    # Border for HUD
    pygame.draw.rect(hud_surface, (100, 200, 255, 200), (0, 0, hud_width, hud_height), 2)

    # HUD Content formatting
    history_str = " -> ".join([NODES[n]["name"] for n in robot.history])
    hud_texts = [
        ("=== ROBOT HUD / INNER VOICE ===", (100, 220, 255)),
        (f"Position: ({int(robot.pos[0])}, {int(robot.pos[1])})", WHITE),
        (f"Direction Vector: ({robot.direction_vector[0]:.2f}, {robot.direction_vector[1]:.2f})", WHITE),
        (f"Next Node: {NODES[robot.next_node]['name']}", WHITE),
        (f"History: [{history_str}]", WHITE),
        (f"Voice: \"{robot.inner_voice}\"", (255, 200, 100))
    ]

    y_offset = 12
    for text, color in hud_texts:
        rendered_text = font.render(text, True, color)
        hud_surface.blit(rendered_text, (12, y_offset))
        y_offset += 26

    # Blit HUD onto main screen at top-left corner
    screen.blit(hud_surface, (5, 5))

    pygame.display.flip()
    clock.tick(30)

pygame.quit()
sys.exit()
 

Die lange Reise zur Künstlichen Intelligenz

 Von der Computertechnik ist bekannt dass sie sich sehr schnell weiterentwickelt. Innovationen wie die 3.5 Zoll Floppy disk waren einerseits Meilensteine des Fortschritts, wurden zugleich aber nach wenigen Jahren durch bessere Techniken wie USB Flash drives ersetzt. Die Computertechnisch schreitet ständig voran.

Ganz anders verlief die Entwicklung der Künstlichen Intelligenz sehr langsam. Die Ursprünge lassen sich auf das Jahr 1912 zurückführen als Torres Quevedo einen Schachautomaten für das Endspiel konstruierte [1] Seite 1. In den 1980er wurden Schachprogramme in Software realisiert und erst 1997 gelang es der Firma IBM unter hohem Technischen Aufwand den menschlichen Weltmeister zu schlagen.[1] seite 2. 

Einfacher formuliert hat es 85 Jahre gedauert bis der Schachautomat von Torres Quevedo soweit verbessert wurde, dass er tatsächlich einsatzfähig war.

Die Langsamkeit der Entwicklung deutet darauf hin, dass die Realisierung Künstlicher Intelligenz anspruchsvoll ist. Trotz hohem Aufwand durch Forscher an den Universitäten und in der Industrie gelang es über Jahrzehnte nicht nennenswerte Fortscrhitte zu erzielen. Gleichzeitig sei erwähnt dass Computerschach ohne praktische Bedeutung ist, will man KI in Form von Robotik einsetzen benötigt man weitere Forschungsprojekte.

Es gibt eine mögliche Erklärung warum die Geschichte der Künstlichen Intelligenz so langatmig ist. Weil ähnlich wie bei dem Versuch ein Perpetuum mobile zu konstruieren die meisten prototypen nicht funktionieren. Von frühen Neuronalen Netzen aus den 1990er Jahren ist bekannt dass sie keinerlei Ergebnis erzielten, es blieb unklar ob das neuranale Netz mittels OCR auf einem Bild eine handschriftliche Zahl erkannte oder nicht. Bei Robotik-Projekte sieht es noch pessimistischer aus. Das Stanford Cart was in den 1970er Jahren von Hans Moravec und anderen entwickelt wurde, war langsam blieb häufig stehen und funktionierte nicht.

Es verwundert wenig das kritischer Beobachter der KI Forschung zu dem Schluss kamen, dass der Ansatz an sich, also einer Maschine das Denken beizubringen, nicht funktioniert.

Quellen:

[1] Bruderer, Herbert. "Die künstliche Intelligenz begann 1912 mit dem Schachautomaten von Torres Quevedo." (2020).