For the specific job of learning to speak, an AI speaking app and a human tutor produce similar scores. A crossover study by Zafer Susoy found no significant difference in speaking achievement between the two. The real differences are anxiety, cost and depth. Use the app to buy daily reps, and a human to buy depth.
The question people are actually asking
Almost nobody types this query because they are curious about pedagogy. They type it because they have been studying a language for years, they can read it, they can follow a film, and then a waiter asks them something and their mouth stops working.
Faced with that, the standard advice is “just get a tutor.” It is good advice that this particular person will not take. They will not book a tutor precisely because a tutor is a witness. Paying a stranger to watch you be bad at something is a hard sell to someone whose entire problem is that being watched is what breaks them.
So invert the question. Instead of asking which option is theoretically superior, ask which one removes the thing that is actually stopping you. That turns out to be a question the research can answer.
What the research says when you put both in the same room
The cleanest test of this comparison is a within-subjects crossover study by Zafer Susoy, published in Frontiers in Psychology. Forty-eight first-year English Language Teaching students at a Turkish public university each sat two comparable speaking exams two weeks apart. One was facilitated by an AI chatbot, the other by human instructors. Both were scored on a TOEFL iBT rubric from 0 to 12, and a speaking anxiety scale was administered before each.
Three results matter.
The scores came out the same. AI condition 8.12, human condition 8.35, t(47) = 1.21, p = 0.23. Not a significant difference. Whatever you think an AI examiner is missing, it did not show up as easier or harder marks. Human and AI ratings also agreed strongly with each other, Cohen’s kappa = 0.86.
Anxiety was lower with the machine, 98.48 against 102.94, p = 0.01, a small effect.
And the finding that reframes the whole comparison: with a human in the room, anxiety predicted the score, r = -0.500, p < 0.01. With the AI, it predicted nothing at all, r = -0.042. The more nervous students lost real marks in front of a person. In front of the machine, their nerves stopped showing up on the scoreboard.
That is not a claim that the AI teaches better. It is a claim that a human audience adds a tax, and that the tax is heaviest on exactly the learners who are already struggling to speak.
Where the human tutor clearly wins
Now the counterweight, because a comparison that only cites the flattering study is an advertisement.
Yi Ma and Mingyang Chen ran a 16 week study with 150 intermediate Chinese learners of English, published in Frontiers in Psychology, with a delayed test at week 20. Three groups: AI with teacher scaffolding, AI alone, and a control.
- IELTS band gains: 1.45 with AI plus a teacher, 0.92 with AI alone, 0.31 for the control.
- Retention at week 20: 94% of the gains held for the scaffolded group, 87% for AI alone.
- The AI-only group hit what the researchers called a novelty plateau around weeks 9 to 12. One learner described being stuck in a loop, with nothing pushing them toward harder material.
- 34% of learners were frustrated by the AI’s lack of cultural adaptability.
- Speech recognition was not equally kind to everyone. Rural learners needed 2.3 times more repetitions than urban peers. One said the AI’s American accent made them feel like a foreigner in their own practice.
Read those numbers honestly. AI alone produced nearly a full IELTS band in 16 weeks, which is real. And the group with a human gained substantially more, plateaued less, and kept more of it. The teacher’s contribution was specific: teacher scaffolding predicted learner autonomy, the sense of directing your own learning, at β = 0.52. The AI’s personalization predicted a different thing, competence, the sense of being good at it, at β = 0.37. They are not substitutes. They feed different needs.
The limitation nobody selling AI wants to print
The chatbot literature is strong on short-term, in-program gains and weak on what happens afterwards.
Mengdi Li, Yinyu Wang and Xiaorong Yang pooled 41 experimental and quasi-experimental studies covering 3,515 participants and found a moderate to large positive effect on second language acquisition, ES = 0.576, 95% CI [0.385, 0.768], p < 0.001. Their own moderator analysis contains the caveat: interventions lasting one to seven days showed among the largest effects. Very short studies producing very large numbers is the classic signature of measuring novelty and immediate practice effects, not durable skill.
There is a second problem, and it is structural rather than statistical. Large language models are trained to be agreeable. Stanford’s SycEval study measured sycophantic behavior in 58.19% of model responses. A conversation partner biased toward accepting what you say is a poor error detector and, worse for this specific job, a poor rehearsal for reality. Real interlocutors interrupt. They get impatient. They do not slow down because you look confused. Practicing only against something infinitely patient leaves one muscle completely untrained.
If you want a single honest sentence on transfer: AI practice reliably builds the reflex of speaking, and human conversation is still where you find out whether the reflex survives contact.
What each one actually costs
Cost is not a detail here. It is the mechanism that decides how much you practice.
Human tutoring, from italki’s own pricing page: community tutors, who are native or fluent speakers without formal teaching training, typically charge between $4 and $20 per lesson. Professional teachers, with qualifications, certifications and structured lesson plans, typically charge between $10 and $40 per lesson. You pay the price on the tutor’s profile, there is no hidden booking fee at checkout, and italki takes its commission from the teacher’s side. Most tutors offer a trial lesson, usually 30 minutes, at 30% to 50% off. Preply is the other large marketplace in this category and works on a subscription of set weekly hours.
That is genuinely reasonable money for expert human attention. But notice what the pricing structure does to behavior. A tutor is a calendar event with a price tag, so you book one or two a week. An AI speaking app is a flat monthly pool of minutes, so the marginal cost of one more ten minute conversation is nothing, and you can have it at 6am or during the walk to the shop. Frequency is where speaking skill actually comes from. The app does not win on quality per hour. It wins on hours.
Who each one is genuinely for
Get a human tutor if: you already speak well enough to hold a conversation and want to go from competent to precise. You need cultural and pragmatic nuance, the register shifts, the things that are technically correct but land wrong. You have a deadline, an exam, a job interview, a move. You know you will not do the work unless someone is expecting you on Tuesday at seven. Or you want somebody who remembers your particular bad habit and refuses to let it slide for the fourth week running.
Use an AI speaking app if: you understand far more than you can produce. You have not said a full sentence out loud in the language in months or years. Your available practice time is fragmented, fifteen minutes here, ten minutes there, at hours when no tutor is awake. You want volume, and you want to make the same mistake forty times without anyone’s face changing.
Use both if you are serious. That is not a diplomatic non-answer, it is the result with the biggest measured gain in the Ma and Chen study. The productive order for most adults is app first, human second: buy the reps and lose the fear with the machine, then bring a much less frightened speaker to a tutor who can finally work on something other than your panic.
Where Mintza fits this argument
Mintza is a live, open voice conversation with a bilingual AI teacher. No transcription-and-wait, no text chat, no lesson script to march through. You talk, it talks back, and it corrects you inside the conversation rather than handing you a report afterwards.
The part that matters for this comparison is the bilingual mechanic. The teacher speaks both the language you are learning and a language you already speak. When you freeze mid-sentence, it drops into the language you already know, unblocks you, and walks you back into the target language. That behavior is enforced by level, not left to chance: at Starter it speaks mostly your own language and introduces the new one in short repeatable chunks, at Beginner it holds most of the conversation in the target language and uses your language to explain and reassure, at Intermediate it stays almost entirely in the target language and drops out only when you are struggling, and at Advanced it stays fully in the target language at natural speed and gives peer-level feedback rather than teacher mode. Correction intensity scales the same way, from modelling with no explicit corrections at Starter, to nuance and register at Advanced.
This is the specific thing that makes speaking from day one survivable. A tutor who happens to share your first language can do it too, and good ones do. But you cannot count on finding one, you pay by the hour for the privilege, and the immersion-only community tutor many learners are pointed toward cannot do it at all. The escape hatch is what stops a beginner from abandoning the sentence, and abandoning the sentence is how people quit.
The rest is built for volume rather than ceremony. Fifteen languages in any pair and either direction. Four levels, and a regional accent picker for the languages that have distinct regional varieties. Both are saved per language. It remembers previous conversations and can pick up where you left off instead of resetting to small talk. No streaks, no clock running on screen while you speak, no imposed path, and reminders only if you set them. You get 20 free minutes when you sign in with Google or Apple, no card, and they never expire. Paid plans are monthly pools of minutes, Basic 180, Plus 360, Pro 600, with no daily cap. For current pricing, check the App Store or Google Play listing, since it varies by country.
The verdict
The honest answer to “app or tutor” is that they are not competing for the same job.
A human tutor is the better teacher. A tutor gives you depth, cultural judgment, accountability and the useful discomfort of a real person who will not accommodate you forever. If you can afford one and you will actually show up, book one. The research supports that: the learners with a human alongside the AI gained the most and kept the most.
An AI speaking app is the better gym. It is where the reps are cheap enough to be daily and the audience is not scoring you while you fumble. The Susoy result is the whole argument in one line: the marks came out the same, and only the human condition charged you for being nervous.
If you understand the language and freeze when someone speaks to you, the sequence is not complicated. Build the reflex where nobody is watching. Then go and be watched.
Related reading: why you can understand a language but can’t speak it, how to practice speaking a language alone, and the best AI apps to practice speaking a language.
Mintza is available for iPhone, iPad and Mac and Android.