
Notice
This article was written based on Turing Test-related papers and research findings publicly available as of August 2, 2026.
The result that “AI passed the Turing Test” may vary depending on the AI model used, the prompt, the length of the conversation, the composition of the evaluators, and the experimental method. Therefore, the results of individual studies should not be interpreted more broadly as meaning that AI possesses the same intelligence or consciousness as humans.
What This Article Covers
When talking with someone online, it can sometimes be difficult to determine whether the other party is a human or an AI.
It may be difficult to tell whether a customer service agent is a real person or an AI chatbot, and it may also be difficult to determine whether comments or posts in an online community were written by a person or created by generative AI.
More recently, as AI voices have become increasingly natural, the same issue has also emerged in telephone and voice conversations.
If AI speaks like a human and people cannot distinguish it from one, can that AI be considered intelligent?
This question did not first emerge in the age of generative AI.
In 1950, British mathematician Alan Turing had already raised a similar question. The evaluation method proposed by Turing later came to be known as the Turing Test.
However, current Turing Test research does not end with simply determining whether AI can converse like a human.
Recent studies have found that AI may be indistinguishable from humans in short text conversations but may find it more difficult to maintain human-like consistency as conversations become longer.
In voice conversations, differences in speaking style and emotional expression have appeared, while group discussions generated to resemble conversations among multiple users have sometimes been difficult to distinguish from real online discussions.
Alan Turing Asked, “Can Machines Think?”
Alan Turing began his 1950 paper, “Computing Machinery and Intelligence,” published in the journal Mind, with the following question.
Can machines think?
However, it is not easy to define precisely what the words “machine” and “think” mean.
Rather than continuing a debate over the definitions of these words, Turing reframed the question in a form that could be observed and tested.
Instead of attempting to determine directly whether thinking was actually taking place inside a machine, he chose to evaluate what kind of behavior the machine displayed externally.
Turing described this as the Imitation Game.
Turing’s paper did not declare that “a machine that cannot be distinguished from a human possesses true intelligence.”
Nor did it propose a method for proving the existence of consciousness or a sense of self.
The key point is that Turing reframed a philosophical question that was difficult to define as a problem of evaluating observable behavior.
How Is the Turing Test Conducted?
The Turing Test, as it is commonly described today, involves three participants.
- A human questioner
- A human respondent
- A machine respondent
The questioner communicates through text with the human respondent and the machine respondent without seeing either of them directly.
The questioner freely asks questions, compares the answers, and then decides which respondent is human and which is a machine.
If the machine converses like a human and the questioner cannot reliably distinguish between the two, the machine is considered to have passed the test.
Why Is the Conversation Conducted Through Text?
Turing proposed delivering answers in writing so that the participants’ voices would not influence the judgment.
The purpose was to distinguish between the human and the machine based only on their answers and manner of conversation, rather than on their voice, face, appearance, or movements.
Therefore, the traditional Turing Test is not an examination of a robot’s appearance or movements.
It is closer to a test that evaluates whether a machine displays behavior similar to that of a human in a natural-language conversation.
Does Being Mistaken for a Human by 30% Mean It Passed?
A commonly cited standard often appears in explanations of the Turing Test.
If 30% or more of the evaluators judge the machine to be human, it has passed the Turing Test.
However, Turing did not define 30% in his paper as an official passing threshold that should apply to every Turing Test.
Turing predicted that, in about 50 years, computers would perform the Imitation Game so well that an average questioner would have no more than a 70% chance of making the correct identification after five minutes of conversation.
Put differently, this would mean that the evaluator could make the wrong judgment about 30% of the time.
However, this was closer to a prediction about the future performance of computers than an official passing rule for every Turing Test.
Even today, there is no single official organization that certifies the Turing Test, nor is there one common test used in every study.
The conditions may differ from one study to another.
- Does the questioner converse with either one human or one AI?
- Does the questioner compare a human and an AI at the same time?
- How many minutes does the conversation last?
- Are the evaluators members of the general public or AI experts?
- What prompt was given to the AI?
- Is only text used?
- Are voice conversations or conversations involving multiple participants also included?
Therefore, when a claim says that “AI passed the Turing Test,” it is necessary to examine the conditions under which the test was conducted.
What Does the Turing Test Evaluate?
The Turing Test primarily evaluates how human-like an AI appears during a conversation.
The evaluation may include the following elements.
- Natural use of language
- Answers that fit the context of the question
- Consistency throughout the conversation
- Human-like speaking styles and expressions
- Emotional and social responses
- The ability to give the other party the impression that it is human
However, the Turing Test alone cannot directly determine the following.
| What the Turing Test Can Evaluate | What the Turing Test Alone Cannot Determine |
|---|---|
| Whether it converses naturally | Whether it actually understands the content |
| Whether it uses a human-like speaking style | Whether its answers are factually correct |
| Whether it maintains the context of the conversation | Whether it has consciousness or a sense of self |
| Whether it responds like a human | Whether it is safe and reliable |
| Whether it makes the other party feel that it is human | Whether it can solve problems in the real world |
Even an incorrect answer can appear human if it is delivered naturally and confidently.
Conversely, even when an AI provides accurate and excellent answers, it may appear machine-generated if its language is overly polished and consistent.
In other words, the ability to appear human and the ability to solve problems accurately are not the same evaluation target.
Can Generative AI Pass the Turing Test?
GPT-4 Was Difficult to Distinguish from a Human in a Short Conversation
Cameron Jones, Ishika Rathi, Sydney Taylor, and Benjamin Bergen presented “People Cannot Distinguish GPT-4 from a Human in a Turing Test” at the 2025 ACM FAccT conference.
The researchers conducted a preregistered randomized controlled experiment.
Participants held a five-minute text conversation with either a human or an AI and then judged whether their conversation partner was a person or an AI.
The results were as follows.
- The actual human was judged to be human 67% of the time.
- GPT-4 was judged to be human 54% of the time.
- ELIZA, an early rule-based chatbot, was judged to be human 22% of the time.
GPT-4 was not evaluated as being more human-like than the actual human.
However, participants were unable to distinguish GPT-4 at a level more accurate than chance. The researchers interpreted this as empirical evidence that modern AI had passed a two-party Turing Test in which the conversation partner was either a human or an AI.
The researchers also found that participants did not distinguish between humans and AI based only on logical reasoning ability.
Stylistic and socioemotional characteristics, such as speaking style, personality, and emotional expression, played an important role in their judgments.
GPT-4.5 Was Selected as the Human More Often Than the Actual Human
Cameron Jones and Benjamin Bergen expanded the experimental method and published “Large Language Models Pass a Standard Three-Party Turing Test” in the Proceedings of the National Academy of Sciences, or PNAS, in 2026.
In this study, the questioner interacted simultaneously with one actual human and one AI.
Participants conversed through two chat windows for five minutes and then selected which participant they believed was the actual human.
The following systems were used in the experiment.
- ELIZA
- GPT-4o
- Llama 3.1 405B
- GPT-4.5
When the AI systems were given a human-like persona prompt designed to help them maintain a human-like persona and speaking style, the results were as follows.
- GPT-4.5 was selected as the human 73% of the time.
- Llama 3.1 405B was selected as the human 56% of the time.
- ELIZA was selected as the human 23% of the time.
- GPT-4o was selected as the human 21% of the time.
GPT-4.5 was selected as the human more often than the actual human it was compared with.
The result for Llama 3.1 405B was not statistically significantly different from that of the human comparison participant.
The researchers presented these findings as the first empirical evidence that an artificial system could pass a standard three-party Turing Test.
However, performance fell substantially when the same models were not given a human-like persona prompt.
Without the prompt, GPT-4.5 was selected as the human 36% of the time, while Llama 3.1 405B was selected as the human 38% of the time.
This shows that the result of the test is strongly affected not only by a model’s basic performance but also by what role and speaking style are assigned to the AI.
Does Passing the Turing Test Mean That AI Possesses Human Intelligence?
The result that GPT-4.5 was selected as the human more often than an actual human does not prove that AI understands and thinks in the same way humans do.
What the study confirmed was that, under specific experimental conditions, AI could be perceived as a more human-like conversation partner than an actual human.
In particular, the fact that the results changed significantly depending on whether a human-like persona prompt was used shows that the Turing Test is influenced not only by a model’s reasoning ability but also by its speaking style, personality, and social expression.
In other words, the Turing Test requires a distinction between the following two abilities.
- The ability to solve a problem accurately
- The ability to appear human to the other party
An AI with more advanced reasoning ability does not necessarily appear more human.
Conversely, an AI that effectively imitates human speaking styles and emotional expressions is not necessarily more accurate or reliable.
How Far Has Current Turing Test Research Progressed?
Maintaining Human-Like Consistency Becomes More Difficult as Conversations Grow Longer
Appearing human in a short conversation and consistently conversing like the same person over a long period are different abilities.
Weiqi Wu, Hongqiu Wu, and Hai Zhao presented “X-TURING: Towards an Enhanced and Efficient Turing Test for Long-Term Dialogue Agents” as an ACL Main Conference Long Paper in 2025.
The researchers pointed out that the traditional Turing Test focuses on short conversations in which one message is exchanged at a time.
In actual conversations, people may send several messages in succession and must remember previous conversations and settings over a long period while responding consistently.
To evaluate this type of long-term dialogue, X-TURING used a method in which multiple messages were exchanged in succession, along with dialogue histories designed to simulate conversations that had continued over a long period.
In the study, GPT-4 had a pass rate of 51.9% when the conversation lasted three turns, but the rate fell to 38.9% when the conversation lasted ten turns.
The researchers explained that AI showed a tendency to find it more difficult to maintain human-like consistency as the conversation became longer.
This shows that an AI’s ability to produce natural responses for a short period and its ability to converse consistently as the same person over a long period may be different.
As a result, the following question is also becoming important in recent research.
How long can AI maintain a conversation in which it cannot be distinguished from a human?
Research Is Evaluating Group Discussions, Not Only a Single AI
The traditional Turing Test focused on a conversation between one human and one machine.
However, multiple users participate simultaneously in conversations on online communities and social media.
Azza Bouleimen, Giordano De Marzo, Taehee Kim, and other researchers first released “The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions” as an arXiv preprint in 2025. The paper was accepted for publication on July 10, 2026, and was published online in Scientific Reports on July 17.
The researchers collected real discussions among human users on Reddit and created group discussions on the same topics using Llama 3 70B and GPT-4o.
When participants were shown actual human discussions alongside AI-generated discussions, the AI-generated discussions were incorrectly judged to be real human discussions 39% of the time.
In particular, discussions generated by Llama 3 were correctly identified as AI-generated only 56% of the time. This was only slightly higher than random guessing.
The study shows that AI can imitate not only the speaking style of one person but also discussions that appear to involve multiple users with different opinions and roles.
This technology may be used to simulate online communities or study the effects of policies and recommendation systems.
However, if multiple AI accounts produce large amounts of content that appear to be discussions among real people, the technology could also be misused to create artificial public opinion or fake social responses.
The scope of Turing Test research is expanding from individual conversations to online groups and social interactions.
However, as of August 2, 2026, the article page in Scientific Reports stated that the manuscript had been made available before final production editing. Therefore, although the paper had been accepted and published online, some wording or content could still be revised during the final editing process.
AI May Appear Human in Text but Still Sound Different in Voice Conversations
Studies have found that AI can be difficult to distinguish from a human in text conversations, but the situation was different in voice interactions.
Xiang Li, Jiabao Gao, and other researchers evaluated Speech-to-Speech systems that directly listen to speech and respond with speech in “Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction,” presented at ICLR 2026.
Based on conversations involving nine recent voice-conversation systems and 28 human speakers, the researchers collected a total of 2,968 judgments from 397 evaluators.
The results showed that none of the voice AI systems included in the experiment passed the Turing Test.
The problem was not simply that the systems failed to understand the meaning of the answers properly.
The researchers divided human-likeness into 18 detailed elements and explained that the main differences between AI and humans appeared in the following areas.
- Speaking style and intonation
- Emotional expression
- Paralinguistic features such as hesitation and breathing
- Consistency of persona and personality within the conversation
This means that even if an answer appears natural in text, machine-like characteristics may become noticeable in emotional expression, intonation, and conversational style when the answer is heard as speech.
The researchers found that current voice AI produced results close to those of humans in logical consistency and dialogue memory, but differences remained in paralinguistic features, emotional expression, and the expression of conversational personas.
The results were also unstable when general AI models were used as judges to distinguish between humans and machines.
This study shows that current voice AI evaluation is moving beyond simply determining whether the speaker is “a human or a machine.”
It is developing toward identifying in detail which characteristics feel different from those of humans and using those findings to determine how voice AI can be improved.
Current Research Analyzes Conditions and Causes Rather Than Simply Pass or Fail
Recent Turing Test research is moving away from determining a single pass-or-fail result.
The main questions in current research are as follows.
1. Under What Conditions Does AI Become Indistinguishable from Humans?
AI may appear human in a short conversation, but its consistency may break down as the conversation becomes longer.
The results also vary significantly depending on whether a human-like persona prompt is provided.
Therefore, it is necessary to examine the conversation length, prompt, comparison target, and evaluation method, rather than looking only at the model name and pass rate.
2. In Which Forms of Information Do Differences Appear?
AI may be difficult to distinguish from humans in text, but intonation, emotion, and conversational rhythm may make it easier to identify in voice interactions.
In the future, multimodal tests that evaluate facial expressions, video, behavior, and physical interaction together may also become important.
3. Can AI Imitate Group Behavior as Well as Individual Behavior?
AI can generate not only one person’s conversation but also online discussions and community conversations involving what appear to be multiple participants.
As a result, the question of distinguishing between humans and AI is expanding beyond determining the authenticity of an individual conversation to determining the authenticity of online public opinion and social interaction.
4. What Characteristics Do People Consider Human-Like?
In recent studies, people did not necessarily use mathematical or logical questions to distinguish AI from humans.
Socioemotional characteristics such as speaking style, emotional expression, humor, hesitation, and conversational attitude sometimes had a greater influence.
This means that the Turing Test is not only a test of AI. It is also a test that shows what people consider to be human-like.
Why Is Modern AI Not Evaluated Using Only the Turing Test?
The Turing Test is meaningful for evaluating the ability to converse in a way that resembles a human.
However, modern AI systems require many abilities beyond conversation.
- Does the AI provide factually correct answers?
- Can it solve mathematical and logical problems?
- Can it write program code?
- Can it understand images and speech?
- Can it use external tools accurately?
- Can it complete multistep tasks?
- Can it appropriately refuse harmful or prohibited requests?
- Does it avoid producing biased results?
- Can it operate reliably over a long period?
These abilities cannot all be measured through a single conversational test.
For this reason, modern AI research uses multiple benchmarks and real-world task evaluations according to the purpose of the evaluation.
| Evaluation Question | Evaluation Target |
|---|---|
| Does it appear human? | Human-likeness |
| Does it solve problems accurately? | Performance and reasoning |
| Are its answers factually correct? | Accuracy and reliability |
| Can it block dangerous requests? | Safety |
| Can it complete real tasks? | Tool use and task execution |
| Can it maintain consistency over a long period? | Memory and persistence |
Among these questions, the Turing Test evaluates whether AI appears human.
It has limitations as a final test for determining every aspect of AI intelligence.
Is Being Indistinguishable from Humans Always a Good Thing?
When Turing proposed the Imitation Game, a machine that could not be distinguished from a human represented the possibilities of future technology.
However, now that generative AI is widely used, the inability to distinguish between humans and AI can create new problems.
- Should users know that they are interacting with AI?
- Should AI-generated posts and comments be labeled?
- Should AI accounts be allowed to participate in online discussions as if they were real people?
- Could an AI that imitates human emotions and personalities gain an excessive level of trust from users?
- Could the ability to deceive humans be used in fraud or social-engineering attacks?
The ability to appear human can help create more natural services.
However, when it is used to conceal the identity of AI or make it appear to be a real person, problems involving transparency and trust arise.
In the past, being indistinguishable from humans was regarded as a technological achievement. Today, however, whether users should be able to distinguish between humans and AI has also become an important subject of research and policy.
Is the Turing Test a Finished Test?
Some generative AI systems have reached a stage where they can pass the Turing Test under limited text-conversation conditions.
GPT-4.5, when given a specific human-like persona, was selected as the human more often than the actual human in a five-minute, three-party Turing Test.
However, the X-TURING study found that GPT-4’s human-likeness evaluation declined as conversations became longer.
In voice conversations, differences from humans appeared in speaking style, emotional expression, and the consistency of the conversational persona.
Meanwhile, group discussions constructed as if multiple AIs were participating sometimes produced content that was difficult to distinguish from conversations in actual online communities.
Therefore, research on the Turing Test is not finished.
However, the question being studied is changing.
The question in the past was as follows.
Can machines converse like humans?
The current question is more specific.
Under what conditions, for how long, and in which forms of information and social situations does AI become indistinguishable from humans?
The Turing Test can no longer be viewed as a single final test that determines every aspect of AI intelligence.
Instead, its role is expanding as an evaluation framework for studying AI’s human-likeness, the distinguishability of humans and machines, people’s trust, and social interaction.
DANA NOTES in One Line
Some generative AI systems can pass the Turing Test in short text conversations, but current research is developing toward evaluating the conditions and characteristics under which humans and AI can be distinguished, rather than simply determining whether the test was passed.

