Vor dem Aufkommen von instruction following robotern und Vision language action modellen gab es die informierte Suche was ein heuristischer Algorithmus darstellt. Verwendet wurde eine feste Kostenfunktion die beim Pfadplanen häufig der Abstand zum Ziel auf der Karte war während bei Computerschach die Kostenfunktion die Bewertungsfunktion der aktuellen Stellung war, also wieviele Figuren eine Seite besitzt und wo diese stehen auf dem Brett.
Informierte Suche mittels Kostenfunktion ist leistungsfähiger als die vorher übliche backtracking suche welche alle möglichkeiten durchprobiert und eine extrem hhohe laufzeit besitzt. Bei der informierten Suche wird nur ein kleiner Teil des gametree berücksichtigt. Es stellt einen Kompromiss da zwischen einem numerischen Solver wie der potential field methode und eine konkrete Anwendungsdomäne wie dem Navigieren eines Roboters in einem Labyrinth.
Informierte Suche hat jedoch ein größes Problem: die Kostenfunktion ist fest in der Software vorgegeben und kann nicht von außen verändert werden. zusätzlich besteht die Kostenfunktion aus einer mathematischen Formel deren Erstellung schwierig ist und die selten für alle Situationen gute Ergebnisse liefert. Beim Schach wird z.B. eine Summe gebildet aus mehreren Einflussfaktoren die unterschiedlich gewichtet werden.
Die neuere KI Forschung ist deshalb von der klassischen informierten Suche abgerückt zugunsten voice control. Voice control erlaubt es während der Laufzeit des Roboters die kostenfunktion von außen zu verändern. Der User entscheidet während der interaktion was das Ziel ist. er kann z.b. sagen "fahre zur ladestation" oder "fahre zum Wegpunkt A". Ein solches Kommando wird in eine Kostenfunktion übersetzt die dann wiederum konkrete Handlungen des Roboters aktiviert. Genauer gesagt versucht der Roboter ähnlich wie der informierten Suche seine Kosten zu minimieren und das Ziel zu erreichen.
Die einfachste Version eines voice control roboters besteht aus einer Menüstruktur wo der Benutzer aus 4 möglichen Befehlen einen Auswählen kann, für jedes Kommando wurde vorher manuell die Kostenfunktion hinterlegt zwischen denen der Benutzer wählen kann. Das entspricht ungefähr den Optionen beim Reinforcement learning. Bei einer komplexeren voice control steuerung kann der Benutzer halbwegs frei ein Kommando eingeben was in eine inviduelle Kostenfunktion übersetzt wird. z.B. durch das bennnen von wegpunkten oder das benennen von aktionsverben. Einfache Kommandos könnten lauten:
1. fahre langsam zu Wegpunkt B
2. fahre schnell zu Wegpunkt C
3. stopp
4. fahre schnell zu Ladestation
Der Benutzer interagiert hier mit dem Roboter mittels einer simplen Sprachgrammatik. Der Parser hat die Aufgabe für jedes mögliche Kommando eine numerische Kostenfunktion zu bestimmen welche sich mittels potential field Suche in Servobefehle übersetzen lässt.
September 22, 2026
Informierte Suche als Vorläufer von sprachgesteuerten Robotern
September 01, 2026
From mathematics to non mathematics in robotics
The classification matrix shows algorithms on two categories: a) mathematics to linguistics and b) batch to interactive processing. Artificial Intelligence in the past was influenced by the bottom left section based on optimization algorithms like PSO and search algorithms like RRT. These algorithms are useless for advanced robot control so there was a need to invent more advanced techniques.
Advanced means that at first the former focus on mathematics was replaced by a linguistics paradigm. Early examples were ontologies, OWL and knowledge graph. And second the former batch oriented paradigm was replaced by interactive systems. An example which combines both is voice control robotics which is based on interactive with a human user and by natural language.
Voice control was popular in the 2010 for example in the MIT forklift robot and has evolved in more recent vision language action models based on neural networks available since 2025. These state of the art robot control algorithms are located top right in the chart.
Let us take a closer look into the figure. Algorithms in the past were designed with a certain purpose. For example simulated annealing allows to find the local minimum for a cost function which is the correct algorithms more most mathematical optimiziation problems. Other concepts like ontologies were designed to capture domain specific knowledge. It allows a computer to access human knowledge.
The problem with these algorithms was, that they are not powerful enough. Its not possible to use them directly for robot control. Its unclear how a certain robot OWL ontology has to look like and an algorithms like potential field has a very long runtime. So there is a need to develop a new sort of algorithm which is located in a different section of the figure.
This missing Quadrant is located on top right in the figure at the interaction of linguistics + interactive. Such kind of algorithms are very powerful and are new developments. ITs possible to use them for robot control. Their inner working is based on linguistics on the one hand that means, domain knowledge isn't stored in numbers but in words, and secondly they are based on external feedback loops realized with interactive control. That means, a human operator gives a textual command to the robot like "move north and stop".
If these algorithms are labeled with a single term it would be "voice control". These algorithms are working very different from classical AI algorithms in the past because there is no mathematical optimiziation problem and there is no semantic network or knowledge graph available anymore. Instead the algorithm acts as a parser. Its an interface between man and machine.
April 19, 2026
Artificial intelligence with oracle turing machines
Classical turing machines are executing algorithms, therefor the artifical intelligence must be located within an algorithm. There is an extensive list available of all possible algorithm but none of them is providing AI.[1]
There are some algorithms available which are mentioned in the context of AI like automated planning, Mathematical optimization and neural networks, but its not possible to take one algorithm from the list and use it for robot control.
What is needed instead is an opposite computional model different from a turing machine called an oracle turing machine. Even if the mathematical background of such a Super Turing machine is very complex, the principle can be explained as a Turing machine which communicates with an external system. This ability to communicates allows to offloadwing Artificial intelligence.
For robotics application, an oracle turing machine is usally implemented as a teleoperated robot. The robot stops in front of an obstacle and asks the oracle what to do next. The oracle is the human operator who decides that the robot needs to move around the obstacle on the left pathway. This command is executed by the robot.
In contrast to a normal turing machine, an oracle turing machine doesn't process an algorithm but it communicates. Communication means to solve problem by asking someone else outside of the own system. The higher instance is better informated about the situation, a human operator is equipped with a powerful vision system and has a lot of knowledge to solve most robotics problems. Such kind of knowledge is hard to program into an algorithm, so the robot needs to ask the operator for help.
There is a detail problem available in oracle turing machines which is the communication protocol. The turing machine and the oracle need to established a shared communication protocol which allows them to receive and submit messages in a language. This language needs to be invented first.
[1] https://en.wikipedia.org/wiki/List_of_algorithms
February 11, 2025
From heuristics to language guided planning
March 25, 2023
Determine prime numbers with a questionnaire
October 08, 2019
Software design for a grasping robot

Programming a pick&place robot is on the first look a problem for Artificial Intelligence. It has to do with creating an algorithm which is capable of learning grasping poses. A closer look into the problem will show, that AI isn't needed in the domain. Instead, it's an engineering project which has to do with programming a simulation.
The mindmap on top of the posting shows a rough concept of the idea. All the terms used in the chart are domain specific. It's an attempt to formalize the grasping workflow. The mindmap can't executed on a computer, but it's part of a software engineering process. The idea is to program a prototype with the Python language, and the mindmap helps to identify subparts of the project.
The chart is not complete, because a real grasping robot system contains of many more requirements and design principles. Even the task of pick&place looks not very complicated, it can be a demanding project to write a simulator for this purpose.
On the other hand the potential benefit is great. If the grasping domain can be realized in software, this is equal to automation. In many segments of economy, the same task is done millions of time. All supermarkets, all container terminals, all warehouses and most agriculture production facilities are confronted with the simple problem of grasping an object and release it at the target location. Today, most of the work is done by humans, and not by robots. The reason is, that reliable grasping robots are difficult to program. The task is not a toy problem which can be realized in 300 lines of code, but it's a large scale software projects which needs a lot of heuristics preprogrammed into the system.
But let us go into the details. The main idea is to focus on a simulator which is realized with the object oriented paradigm. The domain of “robot grasping” is converted into an UML chart which consists of many classes. The classes are used to store information about the events, the grasp pose, the trajectory of the robot arm, the position of the objects on the table, the result of the vision system and the planned high-level actions. Right now, it's unclear how many classes are needed to model the overall domain. I would guess 100 classes are the minimum requirement for this complex domain.
The problem is, that the pick&place task consists of many subproblems. One of them is called inverse kinematics. Inverse kinematics has to do with controlling a robot gripper indirectly. Even if the inverse kinematics problem was solve, lots of other problems are available for example the gripper speed during the grasp-phase or what to do if the robot gripper lost the object during the transit.
The overall grasping robot will fail, if only one of the subproblems isn't handled well enough. So there is need to structure the overall task hierarchical. I think, that for creating the prototype, it make sense to get an overview with the help of a mindmap. This helps to identify subdomains of the grasping pipeline which can be solved separately. The idea is, to handle the robot task similar to the problem of programming an operating system for an IBM PC. The idea is that every detail has to be handled with sourcecode, and if the projects consists of millions of codelines, the resulting software will run great.
From the perspective of Artificial Intelligence this sounds a bit disappointing, because the AI Community is interested in building simple but powerful system. The secret goal is to program in 500 lines of code an Artificial Intelligence which can learn by itself, which makes software engineering obsolete. This kind of vision can't be realized in reality. Robot projects in reality have the tendency to become complicated and looking similar to normal software engineering projects from game development or application programming.

To minimize the failure probability of the project, it's important to define some constraints in advance and make the domain easy to realize. The first question which has to answer is, which kind of hardware layout make sense for a grasping robot. The most reliable layout is a portal crane which is used in the reality by container terminals. The advantage is, that the overall system can transport heavy loads, it was tested many times for real problems and it's conservative by default.
Other example of robot grasping systems for example a robot arm or a delta robot are interesting for research projects but they were not tested in reality. Usually these designs are utilized for exploring new path and open new research fields. This kind of open ended problem is not needed here.
The second issue which can reduced in complexity are cluttered object grasps. For the beginning it's much easier to avoid these requirements and define that the robot has only grasp normal objects which are aligned in advanced. So it's not a universal grasping robot, but a simple container crane who is working with a repetive mode.
The resulting mindmap looks clearly, the amount of open tasks is small and it's possible to program the prototype with a low amount of codelines. It's important to mentioned, that even with the simplications, it won't become a toy problem but a large scale software project. That means, the task of programming a simulator for a portal crane is highly complex.
