In recent years, Artificial Intelligence (AI) has undergone a massive growth spurt. However, not so long ago, Large Language Models (LLMs) such as ChatGPT, Gemini, and Claude were curiosities. You could trick them, confuse them, or make them contradict themselves. Today, they have evolved into versatile companions that can write software, assist scientific research, extract insights from a large set of documents, and offer structured guidance across a wide range of domains. Today’s multi-modal AI systems no longer operate on text alone; they interpret images, analyse audio, generate video, and combine these streams in seamless ways. Language, reasoning, and creativity, capacities we associated with ourselves are now appearing, at least on the surface, in machines.
Tracing the foundations of these AI systems, one can observe that the core idea behind them is not new. Artificial neural networks have existed since the late 20th century, and their conceptual roots go back even further. In 1943, Warren McCulloch and Walter Pitts proposed a simple mathematical model of a neuron. The McCulloch–Pitts neuron takes numerical inputs, multiplies them by adjustable weights, sums the results, and applies a non-linear function to produce an output. This is similar to how one takes input from multiple people and makes a decision if enough people agree on a course of action. Individually, such units are extremely simple. Yet a powerful mathematical insight, known as the universal approximation theorem, shows that networks composed of enough of these simple units can approximate virtually any function connecting input to output. With sufficient scale, they can process remarkably complex patterns.
Published - February 23, 2026 08:30 am IST