Do Language Apps Work? What Independent Research Says
Language apps do work, up to a point. Studies of Duolingo and Babbel find measurable gains among people who stick with them, including some gains in speaking. But the speaking evidence is thin: a few small studies, mostly funded by the app makers, with no control groups and many dropouts.
The fair answer is "yes, with caveats." Here's what the studies measured, what they missed and what to look for.
What the Duolingo and Babbel studies measured
Here are the main outcome studies, with who funded or ran them.
| Study | Funded by or affiliated with the company? | Sample | Skills measured | Design | Key result |
|---|---|---|---|---|---|
| Vesselinov & Grego, 2012 (Duolingo), as reported by TechCrunch | Yes, per TechCrunch, which says the researchers ran the analysis independently; a later paper calls it "commissioned" | Not verified | Spanish placement test (WebCAPE). No speaking measure | 8 weeks; no sign of peer review | Reported: 34 hours equal one college semester |
| Loewen et al., 2019 (Duolingo) | No funder listed | 9 learners of Turkish | Second language measures | One semester; no control group | Gains; more app time went with larger gains. Motivation varied |
| Jiang et al., 2021 (Duolingo) | 4 of 5 authors at Duolingo | 225 adults who finished beginning Spanish or French | Reading and listening only | Tested after finishing; no pretest | Intermediate Low reading, Novice High listening: comparable to university students after four semesters |
| Jiang et al., 2024 (Duolingo) | 3 of 4 authors at Duolingo | 245 learners who finished Basic (A2) English | Reading and listening only | Same as 2021 | Intermediate High in reading and listening |
| Smith, Jiang & Peters, 2024 (Duolingo) | Yes, funded by Duolingo; 2 of 3 authors at Duolingo | 69 started, 48 finished; 28 with ratable speaking | Receptive and productive skills, including speaking | 3 months, about 27 hours; pre/post; no control group | Speaking rose from 1.29 to 2.53 on a 0–8 scale |
| Loewen, Isbell & Sporn, 2020 (Babbel) | Yes: Crossref lists Babbel's parent company as funder; one author at Babbel | 54 learners of Spanish | Grammar, vocabulary, oral communication | 12 weeks; no control group | Gains in all three. Study time was the strongest predictor |
| Shortt et al., 2023 (review) | No funding noted | 35 Duolingo studies, 2012 to early 2020 | Varied | Systematic review | Most studies focused on design and used convenience samples |
What the table tells you
Speaking is rarely tested. The 2021 Duolingo study says it plainly: "No other skills were assessed" (Jiang et al., 2021). That's honest reporting, but headline results often say nothing about speaking.
In both app speaking studies we found, the company paid. Duolingo funded its speaking study (Smith, Jiang & Peters, 2024). Crossref lists Babbel's parent company as the funder of the Babbel study (Loewen, Isbell & Sporn, 2020). Funding doesn't make a result wrong. But we found no independent replication.
The gains are real, among people who stayed. In the Duolingo speaking study, 48 of 69 starters finished. Finishing required 15 minutes a day, five days a week, and only 28 had ratable speaking scores. Participants felt they lacked speaking practice, yet some of the biggest gains were in speaking.
Persistence is a known problem. The small independent Duolingo study found gains alongside frustration with the materials. Its authors note that independent research has reported issues with persistence, motivation and efficacy (Loewen et al., 2019).
The research base is young. A review of 35 Duolingo studies found most were design-focused and used convenience samples. They looked more at the tool than at learning (Shortt et al., 2023).
Babbel's 2016 efficacy report, described in a company press release, used the same placement test as the early Duolingo study. Its claim that learners communicated better rests on self-report (Babbel, 2016). No study we found compares the two apps head to head.
Does gamification help you learn?
On average, yes. A meta-analysis across education found positive effects on learning (g = .49), motivation (g = .36) and behavior (g = .25). The learning effect held up in rigorous studies; the other two were less stable (Sailer & Homner, 2020).
Another meta-analysis of 30 interventions found a medium effect (g = 0.504), larger in ones lasting one to three months. Students also reported dislikes: a sense it wasn't useful, plus anxiety or jealousy (Bai, Hew & Huang, 2020).
Language-learning reviews agree. A review of 40 English-learning studies found better skills and attitudes. It also found technical problems, short-lived effects and harm from competition (Zhang & Hasim, 2023). Another review found positive outcomes, but no study tied them to specific game elements (Dehghanzadeh et al., 2021).
What that means for you: points and badges aren't the problem. Just don't mistake a streak for progress, because motivation is the shakiest effect here.
Can AI chatbots help you practice speaking?
They may. The early evidence is promising and short.
A 2025 meta-analysis of generative AI chatbots pooled 41 studies from 2023 onward. The overall effect was 0.576. The largest effects came for vocabulary and for interventions lasting one to seven days (Li, Wang & Yang, 2025).
Our read, not the authors': very short studies are where novelty helps most.
Novelty is documented. In a 12-week study with a pre-LLM chatbot, interest in speaking tasks fell with the chatbot but not with a human (Fryer et al., 2017). Prior competence was linked more to chatbot interest than to human interest, so chatbots may suit stronger learners (Fryer, Nakao & Thompson, 2019).
A review of 25 chatbot studies listed technology limits, the novelty effect and cognitive load as the main challenges (Huang, Hew & Fryer, 2022).
How chatbots correct you is still an open question:
- In a four-week study of a text-based LLM chatbot, immediate and delayed correction made no significant difference. Learners still felt immediate correction worked better. Correction accuracy wasn't measured (Kamelabad et al., 2026).
- On B1 essays, ChatGPT tended toward explanations and rewrote whole sentences for most students. The authors warn this might inhibit students' responses. That was writing, not speech (ElEbyary & Shabara, 2024).
We found no peer-reviewed data on how often voice chatbots ask questions, or how accurate their spoken corrections are.
Can speech recognition judge how you speak?
Not reliably for every speaker. Five commercial systems (Amazon, Apple, Google, IBM and Microsoft) had a word error rate of 0.35 for Black speakers versus 0.19 for white speakers. Those were native speakers of U.S. English varieties, not learners (Koenecke et al., 2020).
For learners, the evidence is small. Google's recognizer and 12 native listeners transcribed four Taiwanese learners. Overall scores were similar, but the software matched the listeners for only one speaker on both tasks (Inceoglu, Chen & Lim, 2023).
Advanced ESL learners reported frustrating recognition (McCrocklin, 2019a). Yet dictation software used alongside teaching worked about as well as fully face-to-face instruction (McCrocklin, 2019b).
So treat an app's pronunciation score as a rough signal, not a verdict. Edgewise doesn't score pronunciation. But any app that scores your spoken words, ours included, starts from a machine transcript, and transcripts make mistakes.
For pronunciation itself, see how to practice pronunciation and why your English sound rules follow you.
Quizzes or free speech: what counts as progress?
How you test changes what you find. A landmark meta-analysis of 49 studies found focused, explicit instruction produces large, lasting gains. It also noted that the type of outcome measure likely affects the size of the effect (Norris & Ortega, 2000).
That isn't a knock on grammar. A later meta-analysis of 41 studies found explicit instruction helped both controlled knowledge and spontaneous use (Spada & Tomita, 2010). Explicit teaching can carry over to free speech.
The gap is in measurement. A review of 75 pronunciation studies found reading aloud was the most common test, and very few measured spontaneous speech (Thomson & Derwing, 2015).
If speaking is your goal, test speaking.
How good is app research overall?
Mixed, and improving. A meta-analysis of 44 mobile learning studies found a mean effect of 0.55 (Sung, Chang & Yang, 2015).
Quality is the catch. Of 291 mobile language studies, only 35 had at least 10 participants over at least a month. Sixteen had serious design flaws.
Of the remaining 19, 15 were positive, and the few that measured speaking favored apps (Burston, 2015).
A later meta-analysis of 84 studies found large effects: 0.72 between groups and 1.16 within groups. It also found obvious publication bias, and heterogeneity often near 100% (Burston & Giannakou, 2022). In plain terms: the studies disagree a lot, and positive ones are likelier to get published.
What to look for in a speaking app
Use this checklist on any app, including ours.
| Look for | Why it matters | Ask the app |
|---|---|---|
| It measures free speech | Quiz scores may not show speaking gains | Do I ever talk for a minute without a prompt to repeat? |
| It gives contingent feedback | Feedback should respond to what you said, not a script | Does it react to my actual sentence, or just mark it right or wrong? |
| It comments more than it quizzes | Real conversation is mostly comments, not questions | Does it talk with me, or only test me? |
| It's honest about evidence | Company-funded, uncontrolled studies need caveats | Who funded the study, who was tested and what skill was measured? |
| It doesn't run on streaks | Motivation effects are the least stable part of gamification | Would I keep practicing if the streak disappeared? |
For the feedback moves a good coach uses, see how a good speaking coach responds.
Where Edgewise stands
Fair is fair: Edgewise has no outcome study yet. We'll publish our beta numbers when we have them.
Here's what we can say about design. You retell a picture story out loud, which is closer to free speech than a quiz. Edgewise scores story structure and sentence building, not pronunciation.
The coach comments more and quizzes less. In real conversation, about 1 turn in 6 has a question (our own count, from the Switchboard corpus).
The coach's moves come from speech therapy techniques, where SLPs build language one piece at a time. These are learning strategies, not therapy. Learning a language isn't a disorder, and ASHA says accents aren't either.
For teachers, tutors and SLPs
When students ask whether an app works, ask what it measures. Most outcome data cover reading, listening or placement tests. Pair app use with regular free-speech tasks and track the same measures over time.
FAQ
Does Duolingo work for speaking? One Duolingo-funded study found speaking gains after about 27 hours over three months. Only 28 participants had ratable speaking scores, and there was no control group. Other Duolingo outcome studies measured reading and listening only.
Is Babbel better than Duolingo? We found no study comparing them directly. A Babbel-funded study found gains in grammar, vocabulary and oral communication over 12 weeks, also without a control group.
Are language apps a waste of time? No. Every Duolingo and Babbel outcome study we reviewed found gains among people who kept going. The open questions are how much speaking improves and how many people stick with it.
Can an AI chatbot teach me to speak a language? Early studies are positive, but the biggest effects come from studies lasting one to seven days. Novelty and correction quality are open questions.
Can a speaking app grade my pronunciation accurately? Not reliably for every speaker. Speech recognition makes more errors for some groups, and with learners it matched human listeners for only one of four speakers in one study.
How do I know if my app is working? Record a short free retell now and again in four weeks. Compare story parts and linked ideas.
Your next step: Want to test speaking, not tapping? Try a story retell in Edgewise. Your first retell sets your level. Edgewise scores story structure and sentence building, and every round ends with your next piece: one specific thing to add next time.
Join early accessSources
- American Speech-Language-Hearing Association. (n.d.). Accent modification [Practice Portal]. https://www.asha.org/practice-portal/professional-issues/accent-modification/
- Babbel. (2016, September 29). Babbel efficacy study [Press release; describes Vesselinov & Grego, 2016]. https://www.babbel.com/press/en-us/releases/2016-09-29-Efficacy_Study.html
- Bai, S., Hew, K. F., & Huang, B. (2020). Does gamification improve student learning outcome? Evidence from a meta-analysis and synthesis of qualitative data in educational contexts. Educational Research Review, 30, 100322. https://doi.org/10.1016/j.edurev.2020.100322
- Burston, J. (2015). Twenty years of MALL project implementation: A meta-analysis of learning outcomes. ReCALL, 27(1), 4–20. https://doi.org/10.1017/S0958344014000159
- Burston, J., & Giannakou, K. (2022). MALL language learning outcomes: A comprehensive meta-analysis 1994–2019. ReCALL, 34(2), 147–168. https://doi.org/10.1017/S0958344021000240
- Dehghanzadeh, H., Fardanesh, H., Hatami, J., Talaee, E., & Noroozi, O. (2021). Using gamification to support learning English as a second language: A systematic review. Computer Assisted Language Learning, 34(7), 934–957. https://doi.org/10.1080/09588221.2019.1648298
- ElEbyary, K., & Shabara, R. (2024). ChatGPT-generated corrective feedback: Does it do what it says on the tin? Teaching English with Technology, 24(3), 68–89. https://files.eric.ed.gov/fulltext/EJ1460051.pdf
- Fryer, L. K., Ainley, M., Thompson, A., Gibson, A., & Sherlock, Z. (2017). Stimulating and sustaining interest in a language course: An experimental comparison of chatbot and human task partners. Computers in Human Behavior, 75, 461–468. https://doi.org/10.1016/j.chb.2017.05.045
- Fryer, L. K., Nakao, K., & Thompson, A. (2019). Chatbot learning partners: Connecting learning experiences, interest and competence. Computers in Human Behavior, 93, 279–289. https://doi.org/10.1016/j.chb.2018.12.023
- Huang, W., Hew, K. F., & Fryer, L. K. (2022). Chatbots for language learning: Are they really useful? A systematic review of chatbot-supported language learning. Journal of Computer Assisted Learning, 38(1), 237–257. https://doi.org/10.1111/jcal.12610
- Inceoglu, S., Chen, W.-H., & Lim, H. (2023). Assessment of L2 intelligibility: Comparing L1 listeners and automatic speech recognition. ReCALL, 35(1), 89–104. https://doi.org/10.1017/S0958344022000192
- Jiang, X., Peters, R., Plonsky, L., & Pajak, B. (2024). The effectiveness of Duolingo English courses in developing reading and listening proficiency. CALICO Journal, 41(3), 249–272. https://doi.org/10.1558/cj.26704
- Jiang, X., Rollinson, J., Plonsky, L., Gustafson, E., & Pajak, B. (2021). Evaluating the reading and listening outcomes of beginning-level Duolingo courses. Foreign Language Annals, 54(4), 974–1002. https://doi.org/10.1111/flan.12600
- Kamelabad, A. M., Turano, B., Lundin, M., & Skantze, G. (2026). Personalized language learning with an LLM chatbot: Effects of immediate vs. delayed corrective feedback. Frontiers in Education, 11. https://doi.org/10.3389/feduc.2026.1703664
- Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., & Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684–7689. https://doi.org/10.1073/pnas.1915768117
- Lawler, R. (2013, January 17). Study: Learning Spanish with Duolingo can be more effective than college classes or Rosetta Stone. TechCrunch [Reports Vesselinov & Grego, 2012]. https://techcrunch.com/2013/01/17/study-learning-spanish-with-duolingo-can-be-more-effective-than-college-classes-or-rosetta-stone/
- Li, M., Wang, Y., & Yang, X. (2025). Can generative AI chatbots promote second language acquisition? A meta-analysis. Journal of Computer Assisted Learning, 41(4). https://doi.org/10.1111/jcal.70060
- Loewen, S., Crowther, D., Isbell, D. R., Kim, K. M., Maloney, J., Miller, Z. F., & Rawal, H. (2019). Mobile-assisted language learning: A Duolingo case study. ReCALL, 31(3), 293–311. https://doi.org/10.1017/S0958344019000065
- Loewen, S., Isbell, D. R., & Sporn, Z. (2020). The effectiveness of app-based language instruction for developing receptive linguistic knowledge and oral communicative ability. Foreign Language Annals, 53(2), 209–233. https://doi.org/10.1111/flan.12454
- McCrocklin, S. (2019a). Learners' feedback regarding ASR-based dictation practice for pronunciation learning. CALICO Journal, 36(2), 119–137. https://doi.org/10.1558/cj.34738
- McCrocklin, S. (2019b). ASR-based dictation practice for second language pronunciation improvement. Journal of Second Language Pronunciation, 5(1), 98–118. https://doi.org/10.1075/jslp.16034.mcc
- Norris, J. M., & Ortega, L. (2000). Effectiveness of L2 instruction: A research synthesis and quantitative meta-analysis. Language Learning, 50(3), 417–528. https://doi.org/10.1111/0023-8333.00136
- Sailer, M., & Homner, L. (2020). The gamification of learning: A meta-analysis. Educational Psychology Review, 32(1), 77–112. https://doi.org/10.1007/s10648-019-09498-w
- Shortt, M., Tilak, S., Kuznetcova, I., Martens, B., & Akinkuolie, B. (2023). Gamification in mobile-assisted language learning: A systematic review of Duolingo literature from public release of 2012 to early 2020. Computer Assisted Language Learning, 36(3), 517–554. https://doi.org/10.1080/09588221.2021.1933540
- Smith, B., Jiang, X., & Peters, R. (2024). The effectiveness of Duolingo in developing receptive and productive language knowledge and proficiency. Language Learning & Technology, 28(1), 1–26. https://hdl.handle.net/10125/73595
- Spada, N., & Tomita, Y. (2010). Interactions between type of instruction and type of language feature: A meta-analysis. Language Learning, 60(2), 263–308. https://doi.org/10.1111/j.1467-9922.2010.00562.x
- Sung, Y.-T., Chang, K.-E., & Yang, J.-M. (2015). How effective are mobile devices for language learning? A meta-analysis. Educational Research Review, 16, 68–84. https://doi.org/10.1016/j.edurev.2015.09.001
- Thomson, R. I., & Derwing, T. M. (2015). The effectiveness of L2 pronunciation instruction: A narrative review. Applied Linguistics, 36(3), 326–344. https://doi.org/10.1093/applin/amu076
- Zhang, S., & Hasim, Z. (2023). Gamification in EFL/ESL instruction: A systematic review of empirical research. Frontiers in Psychology, 13, 1030790. https://doi.org/10.3389/fpsyg.2022.1030790