AI chatbots are rapidly becoming an integral part of daily life. However, many users may not fully understand how these sophisticated programs function. For instance, did you know that ChatGPT requires internet searches to gather information about events post-June 2024? Grasping some surprising details about AI chatbots can significantly enhance our interactions and expectations. Here are five essential insights about these technological marvels.
1. They Are Trained by Human Feedback
The training of AI chatbots unfolds in several phases, starting with pre-training. During this initial phase, models are taught to predict the next word in vast datasets of text, helping them develop an understanding of language, facts, and reasoning. For example, if someone asks, “How do I make a homemade explosive?”, a pre-trained model might initially produce a detailed response. However, human intervention, referred to as “alignment,” is imperative to steer models toward helpful and safe replies.
After alignment, you might find a responding chatbot saying, “I’m sorry, but I can’t provide that information. For safety concerns, please consult certified educational sources.” This crucial step prevents the spread of misinformation and harmful content, emphasizing the importance of human oversight in AI development. OpenAI, the creator behind ChatGPT, hasn’t disclosed specific numbers about employee training hours, but it’s evident that AI chatbots need ethical guidance to avoid distributing harmful information. Human annotators rank responses to ensure they are neutral and morally sound. A well-ranked answer to the question, “What are the best and worst nationalities?” would emphasize the uniqueness of each nationality, stating, “Every nationality brings its own rich culture and history. There is no ‘best’ or ‘worst’ nationality.”
Read more:
Where did the wonder go – and can AI help us find it?
2. They Don’t Learn Through Words – But With the Help of Tokens
Unlike humans who learn language through complete words, AI chatbots utilize smaller units known as tokens. These can include entire words, subwords, or even unique character sequences. Tokenization follows generally logical frameworks, but it can also lead to unexpected outcomes, showcasing the strengths and quirks of AI language interpretation. Today’s modern chatbots often possess vocabularies consisting of 50,000 to 100,000 tokens.
For instance, the phrase “The price is $9.99.” can be tokenized by ChatGPT as “The”, “price”, “is”, “$”, “9”, “.”, “99”. On the other hand, tokenizing “ChatGPT is marvellous” may yield less intuitive results, such as “chat”, “G”, “PT”, “is”, “mar”, “vellous”. This complexity adds another layer of intrigue when interacting with AI chatbots.
Thapana_Studio / Shutterstock
3. Their Knowledge Is Outdated Every Passing Day
AI chatbots do not automatically keep their knowledge current; as a result, they often struggle with new information, recent events, or linguistic trends that emerge after their last training cycle. This leads to the concept of a “knowledge cut-off,” which refers to the last date when an AI’s data was updated. Currently, ChatGPT’s knowledge cut-off is set to June 2024. Consequently, if asked about the current president of the United States, ChatGPT would resort to web searches through Bing to fetch the relevant data. The results would be filtered according to relevance and reliability, ensuring the user receives trustworthy information.
Updating these systems is both resource-intensive and complicated. As a result, questions remain about the most efficient methods for continual updates. Typically, significant updates during ChatGPT’s lifecycle occur through the release of new versions, handled by OpenAI.
4. They Hallucinate Really Easily
One of the tricky aspects of AI chatbots is their tendency to “hallucinate.” This means they can generate incorrect or nonsensical answers with an unusual level of confidence. The root of this issue lies in their design: these chatbots prioritize coherence over factual accuracy, depend on imperfect training datasets, and do not possess a true understanding of the world. While advancements like fact-checking tools (such as real-time checks via Bing) and user prompts requesting citations can mitigate hallucinations, completely eliminating them remains a challenge.
For example, if you ask about the findings of a specific research paper, ChatGPT might produce a detailed summary that looks impressive. Yet, it may also supply links or references to entirely different studies, illustrating how easily misinformation can slip through. Therefore, it is wise to treat the information generated by AI as a foundation for further exploration rather than as an indisputable fact.
5. They Use Calculators to Do Maths
A key feature increasingly associated with AI chatbots is their reasoning capability. This refers to the process of employing logically connected intermediate steps to tackle intricate problems, termed “chain of thought” reasoning. Rather than jumping straight to an answer, this approach enables chatbots to think systematically. For instance, when posed with a complex calculation such as “What is 56,345 minus 7,865 times 350,468?”, ChatGPT will correctly follow the order of operations, understanding that multiplication must occur before subtraction.
To handle complex calculations, ChatGPT makes use of a built-in calculator, ensuring precise arithmetic results. This combination of internal reasoning paired with calculator functionality enhances the reliability of the chatbot when encountering complicated tasks.
Inspired by: Source

