How to Practice Pronunciation: Motor Learning Principles for Language Learners
To improve your pronunciation, practice it like a motor skill. Speech-language pathologists use the principles of motor learning (Maas et al., 2008): how much, how often, how varied and in what order you practice, plus what kind of feedback you get, how often and when. Practice that feels harder today often sticks better tomorrow.
These are training principles borrowed from speech therapy, not therapy. Your goal is to be understood with ease and to speak with confidence. It is not to erase your accent.
Pronunciation is partly a motor skill
Mastering the sounds of a second language involves a component of speech-motor skill. That's the premise of a study that applied one motor learning principle, practice variability, to adult learners. With eight Korean adults practicing the English r, varied practice beat practicing a single word in the short term. Long-term learning was limited for everyone (Bu et al., 2021). The evidence in second language learners is early. The principles are well tested in other motor skills, while direct tests in speech are fewer (Maas et al., 2008).
Pronunciation instruction works in general. A meta-analysis of 86 studies found a large average effect, with larger effects for longer interventions and for training that included feedback (Lee, Jang & Plonsky, 2015). A review of 75 studies found that very few measured spontaneous speech (Thomson & Derwing, 2015). So test yourself in real talk, not just in drills.
Performance is not learning
Motor learning research separates two things:
- Performance: how well you do during practice.
- Learning: what you keep after a break (retention) and use in new words and contexts (transfer).
They can point in opposite directions. In a classic study, people who practiced three motor tasks in random order showed better transfer than people who practiced them in blocks, and better retention when test conditions changed (Shea & Morgan, 1979). Random practice often feels worse at the time. A 2026 review describes older adults learning a voice technique: blocked practice best supported acquisition, while random practice tended to improve retention and transfer despite poorer immediate performance (Madill & Ballard, 2026). Speech sounds show a similar pattern. Speakers who trained on nonwords with dissimilar sounds made more errors in training, then were more accurate on new nonwords (Meigh & Kee, 2020).
If practice feels too smooth, it may be too easy.
Step 1: Find the sound (pre-practice)
Before real practice, you need a few correct tries. Clinicians call this pre-practice. Feedback here is typically frequent and immediate, and it focuses on how you made the sound (Madill & Ballard, 2026).
- Hear it. Make sure you can tell the target from your English habit. If you can't, start with ear training.
- Place it. Use one clear placement cue, such as "round your lips" or "no puff of air."
- Get three good tries. Then move on. Pre-practice is short.
Step 2: Set the practice conditions
Maas et al. (2008) review practice and feedback conditions that shape motor learning. Here are six practice levers, adapted for pronunciation.
| Lever | What it means | Early on | Later | Source |
|---|---|---|---|---|
| Amount | How many good tries you get | Many short reps | Still many, spread over days | ASHA, summarizing Maas et al. |
| Distribution | Spacing across sessions | Short daily blocks | Keep sessions spaced | Madill & Ballard, 2026 |
| Variability | Different words, positions, speeds | One or two words | Many words and sentences | Bu et al., 2021 |
| Schedule | Blocked (same target in a row) vs random (mixed) | Blocked | Random | Shea & Morgan, 1979 |
| Attentional focus | Internal (your body) vs external (the effect) | Placement cues help | Focus on the sound and the meaning | Wulf, 2013 |
| Target complexity | How hard the unit is | Syllables, words | Phrases, sentences, retells | ASHA |
A few notes on the levers:
- Distribution and variability. The same review, which applies these principles to voice therapy, notes that spaced sessions support retention better than massed practice, and varied practice promotes transfer (Madill & Ballard, 2026).
- Attentional focus. In motor learning studies, an external focus on the movement's effect tends to beat an internal focus on body movements, for both performance and learning (Wulf, 2013). For pronunciation, "make it sound like sous" is an external focus. "Round your lips" is internal. Use the second to find the sound. Use the first to keep it. In voice training, internally focused cues improved immediate accuracy but didn't consistently improve retention (Madill & Ballard, 2026).
- Target complexity. One SLP approach, speech motor chaining, moves practice from syllables to words, longer words, phrases and self-generated sentences (ASHA). Borrow that ladder, and put a retell on the top rung.
Step 3: Set the feedback conditions
There are two main kinds of feedback:
- Knowledge of performance (KP): feedback about how you made the sound. "Your lips weren't rounded."
- Knowledge of results (KR): feedback about whether it was correct. "Yes" or "not yet."
As practice moves on, feedback should shift from frequent, immediate KP to less frequent, slightly delayed KR (Madill & Ballard, 2026). The same review is blunt: "therapy should not remain in blocked, high-feedback practice." Neither should your self-study.
Two caveats keep this honest. First, feedback frequency is not magic. In one study, 32 English speakers learning new nonwords, some with unfamiliar consonant clusters, all improved, whether they got feedback on 100%, 50%, 20% or 0% of trials (Lowe & Buchwald, 2017). Second, feedback still matters. Japanese learners of English who got recasts during focused practice on the English r improved, even in spontaneous speech. Peers without feedback did not show significant change (Saito & Lyster, 2012). For how recasts and prompts work in conversation, see speech therapy techniques for language learners.
Visual feedback counts too. English-speaking learners of Spanish who trained with visual feedback on voice onset time improved, and the gains carried over to more continuous and spontaneous speech (Offerman & Olson, 2016).
Five targets, five practice plans
Spanish tap and trill: pero vs. perro
Spanish contrasts a tap and a trill between vowels: pero [ˈpe.ɾo] "but" vs. perro [ˈpe.ro] "dog," and caro [ˈka.ɾo] "expensive" vs. carro [ˈka.ro] "car." At the start of a word, Spanish uses the trill: rosa [ˈro.sa].
- Tap. You may already make it. The American English sound in the middle of butter or ladder is very similar to the Spanish tap (Daidone & Darcy, 2014).
- Trill. This one is physics. The tap is a single, quick tongue movement. In the trill, air pressure builds behind the tongue tip and pushes it away from the ridge behind your teeth, again and again, usually about three contacts. Tongue tension and airflow have to balance (Gibson, 2015). Don't flick faster. Set the tip near the ridge, keep it steady but not stiff, and let the air do the work.
- Plan. Pre-practice para until it's a single tap. Block ten tries of perro. Then mix pero, perro, caro, carro, rosa in random order, checking only every other try.
Unaspirated p, t, k: Spanish, French and Portuguese
English puts a puff of air after p, t, k at the start of a stressed syllable. Spanish doesn't. The difference shows up in voice onset time (VOT), the time between releasing the consonant and the start of voicing (Haskins Laboratories). English word-initial voiceless stops fall roughly between 30 and 120 milliseconds. Spanish /p t k/ fall between 0 and 30 milliseconds (Amengual, 2023). French /p t k/ are unaspirated too (Fougeron & Smith, 1993), as are Brazilian Portuguese ones (Kupske & de Oliveira, 2020).
- Find it. English already has the target after s. Compare pin and spin, or key and ski (Broeders & Gussenhoven). Say spa, then drop the s. That's Spanish pa.
- Feedback. Hold a tissue or the back of your hand near your lips. No flutter, no puff. That's KR you can give yourself.
- Why it happens. Aspiration is an English rule you run without noticing. See why your English sound rules follow you.
- Words. Spanish papá [paˈpa], taco [ˈt̪a.ko], pato [ˈpa.t̪o]. The little bridge under [t̪] means the tongue touches the back of the teeth, which is how Spanish makes its t. French tout [tu], tu [ty]. Portuguese pau [ˈpaw].
French u vs. ou, and nasal vowels
French separates su [sy] "known" from sous [su] "under," tu [ty] "you" from tout [tu] "all," and rue [ʁy] "street" from roue [ʁu] "wheel." French [y] has the tongue position of [i] with rounded lips.
- Find it. Say ee and hold your tongue still. Round your lips. That's [y].
- Don't skip ou. In a classic study, English-speaking learners of French produced the new vowel [y] more accurately than [u], which feels familiar but differs from English (Flege, 1987, as summarized by Oakley, 2019).
- Nasal vowels. Contrast beau [bo] with bon [bɔ̃], lent [lɑ̃] with long [lɔ̃], and vin [vɛ̃], vent [vɑ̃], vont [vɔ̃]. Keep your tongue off the roof of your mouth at the end. No n.
Korean lax, tense and aspirated stops
Korean has three p-like stops, three t-like stops and three k-like stops: 불 [pul] "fire," 뿔 [p͈ul] "horn," 풀 [pʰul] "grass." Same for 달 [tal] "moon," 딸 [t͈al] "daughter," 탈 [tʰal] "mask."
Two cues carry the contrast. Aspirated stops have the longest puff, lax stops a medium one, tense stops the shortest. Tense and aspirated stops start the vowel on a high pitch. Lax stops start it low (Choi, Kim & Cho, 2020). In Seoul Korean, pitch is taking over from the puff as the main cue between lax and aspirated stops, especially at the start of a phrase (Silva, 2006; Choi, Kim & Cho, 2020).
- Plan. Practice pitch as part of the sound. Lax: start low. Aspirated: puff and high. Tense: no puff, high and tight. Run all three words in random order once each is stable.
Portuguese nasal vowels and diphthongs
Portuguese contrasts oral and nasal diphthongs (transcriptions here are Brazilian): pau [ˈpaw] "stick" vs. pão [ˈpɐ̃w̃] "bread," and mau [ˈmaw] "bad" vs. mão [ˈmɐ̃w̃] "hand." Add mãe [ˈmɐ̃j̃] "mother."
- Find it. Let air flow through your nose and mouth together for the whole vowel and glide. Check it: pinch your nose and say pau. It shouldn't change. Say pão. It should sound blocked.
- Plan. Block pão until it's stable. Then mix pau, pão, mau, mão, mãe. Then say them in phrases: o pão da mãe.
FAQ
Is pronunciation really a motor skill? Partly. Researchers describe second language pronunciation as involving a component of speech-motor skill (Bu et al., 2021). Perception and knowledge of the sound system matter too, which is why ear training belongs in your plan.
Should I drill one sound at a time or mix them? Both, in order. Block practice on one target helps you find the sound. Random practice across targets is harder during practice but tends to produce better retention and transfer.
How much feedback should I get? A lot at first, focused on how you made the sound. Then fade it to occasional right-or-wrong feedback, slightly delayed, so you learn to judge your own speech.
Do I need to sound like a native speaker? No. The goal is intelligibility and confidence. Research on second language speech shows that a strong accent does not necessarily make speech harder to understand (Munro & Derwing, 1995).
Can I give myself feedback without a teacher? Yes. Record and compare with a model, use a tissue test for aspiration, or use visual feedback tools. Studies with visual feedback show gains in learners' sounds (Kartushina et al., 2015; Offerman & Olson, 2016).
Why is the Spanish rolled r so hard for English speakers? Most English accents have no trill, and the trill is driven by airflow, not by flicking the tongue. English-speaking learners of Spanish can usually hear the tap–trill difference but find it hard to produce (Daidone & Darcy, 2014). Treat the trill as a motor skill: a steady tongue tip, airflow, and many spaced tries. Start with the tap, which American English speakers often already have.
Your next step: Drills build the sound. Retells test it. Try a story retell in Edgewise: real connected speech, where your target words have to hold up. Edgewise scores story structure and sentence building, not pronunciation, and ends every round with your next piece to add.
Join early accessSources
- Amengual, M. (2023). The acoustic realization of L2 Spanish phonetic categories and allophonic alternations by English-speaking immigrants. Proceedings of the 20th International Congress of Phonetic Sciences (ICPhS 2023, Prague), paper 154. https://www.internationalphoneticassociation.org/icphs-proceedings/ICPhS2023/full_papers/154.pdf
- American Speech-Language-Hearing Association. (n.d.). Childhood apraxia of speech [Practice Portal; summary of principles of motor learning]. https://www.asha.org/practice-portal/clinical-topics/childhood-apraxia-of-speech/
- American Speech-Language-Hearing Association. (n.d.). Speech sound disorders: Articulation and phonology [Practice Portal]. https://www.asha.org/practice-portal/clinical-topics/articulation-and-phonology/
- Broeders, T., & Gussenhoven, C. (n.d.). 9.2 Aspiration. In An introduction to American English phonetics. University of Groningen Open Textbooks. https://opentextbooks.rug.nl/americanenglishphonetics2/chapter/9-2-aspiration/
- Bu, L., Nagano, M., Harel, D., & McAllister, T. (2021). Effects of practice variability on second-language speech production training. Folia Phoniatrica et Logopaedica, 73(5), 384–400. https://doi.org/10.1159/000510621
- Choi, J., Kim, S., & Cho, T. (2020). An apparent-time study of an ongoing sound change in Seoul Korean: A prosodic account. PLOS ONE, 15(10), e0240682. https://doi.org/10.1371/journal.pone.0240682
- Daidone, D., & Darcy, I. (2014). Quierro comprar una guitara: Lexical encoding of the tap and trill by L2 learners of Spanish. In R. T. Miller et al. (Eds.), Selected proceedings of the 2012 Second Language Research Forum (pp. 39–50). Cascadilla Proceedings Project. https://www.lingref.com/cpp/slrf/2012/paper3084.pdf
- Flege, J. E. (1987). The production of "new" and "similar" phones in a foreign language: Evidence for the effect of equivalence classification. Journal of Phonetics, 15(1), 47–65. https://doi.org/10.1016/S0095-4470%2819%2930537-6
- Fougeron, C., & Smith, C. L. (1993). French. Journal of the International Phonetic Association, 23(2), 73–76. https://doi.org/10.1017/S0025100300004874
- Gibson, M. (2015). A stochastic approach to rhotic variation in Spanish codas. Loquens, 2(1), e015. https://doi.org/10.3989/loquens.2015.015
- Haskins Laboratories. (n.d.). Legacy: Abramson/Lisker VOT stimuli. https://www.haskinslaboratories.org/vot
- Kartushina, N., Hervais-Adelman, A., Frauenfelder, U. H., & Golestani, N. (2015). The effect of phonetic production training with visual feedback on the perception and production of foreign speech sounds. The Journal of the Acoustical Society of America, 138(2), 817–832. https://doi.org/10.1121/1.4926561
- Kupske, F. F., & De Oliveira, M. S. (2020). O desenvolvimento do padrão de voice onset time das oclusivas surdas iniciais do inglês por aprendizes soteropolitanos: Efeitos da instrução explícita [In Portuguese]. Ilha do Desterro, 73(3), 185–204. https://doi.org/10.5007/2175-8026.2020v73n3p185
- Lee, J., Jang, J., & Plonsky, L. (2015). The effectiveness of second language pronunciation instruction: A meta-analysis. Applied Linguistics, 36(3), 345–366. https://doi.org/10.1093/applin/amu040
- Lowe, M. S., & Buchwald, A. (2017). The impact of feedback frequency on performance in a novel speech motor learning task. Journal of Speech, Language, and Hearing Research, 60(6S), 1712–1725. https://doi.org/10.1044/2017_JSLHR-S-16-0207
- Maas, E., Robin, D. A., Austermann Hula, S. N., Freedman, S. E., Wulf, G., Ballard, K. J., & Schmidt, R. A. (2008). Principles of motor learning in treatment of motor speech disorders. American Journal of Speech-Language Pathology, 17(3), 277–298. https://doi.org/10.1044/1058-0360%282008/025%29
- Madill, C., & Ballard, K. (2026). Optimizing voice therapy interventions: The application of the principles of motor learning in clinical practice. Current Opinion in Otolaryngology & Head & Neck Surgery. Advance online publication; open access at PMC13152083. https://doi.org/10.1097/MOO.0000000000001126
- Meigh, K. M., & Kee, E. (2020). Dissimilar phonemes create a contextual interference effect during a nonword repetition task. Frontiers in Psychology, 11, 585745. https://doi.org/10.3389/fpsyg.2020.585745
- Munro, M. J., & Derwing, T. M. (1995). Foreign accent, comprehensibility, and intelligibility in the speech of second language learners. Language Learning, 45(1), 73–97. https://doi.org/10.1111/j.1467-1770.1995.tb00963.x
- Oakley, M. (2019). Articulation of L2 French mid and high vowels. Proceedings of the 19th International Congress of Phonetic Sciences (ICPhS 2019, Melbourne). https://assta.org/proceedings/ICPhS2019Microsite/pdf/full-paper_414.pdf
- Offerman, H. M., & Olson, D. J. (2016). Visual feedback and second language segmental production: The generalizability of pronunciation gains. System, 59, 45–60. https://doi.org/10.1016/j.system.2016.03.003
- Saito, K., & Lyster, R. (2012). Effects of form-focused instruction and corrective feedback on L2 pronunciation development of /ɹ/ by Japanese learners of English. Language Learning, 62(2), 595–633. https://doi.org/10.1111/j.1467-9922.2011.00639.x
- Shea, J. B., & Morgan, R. L. (1979). Contextual interference effects on the acquisition, retention, and transfer of a motor skill. Journal of Experimental Psychology: Human Learning and Memory, 5(2), 179–187. https://doi.org/10.1037/0278-7393.5.2.179
- Silva, D. J. (2006). Acoustic evidence for the emergence of tonal contrast in contemporary Korean. Phonology, 23(2), 287–308. https://doi.org/10.1017/S0952675706000911
- Thomson, R. I., & Derwing, T. M. (2015). The effectiveness of L2 pronunciation instruction: A narrative review. Applied Linguistics, 36(3), 326–344. https://doi.org/10.1093/applin/amu076
- Wulf, G. (2013). Attentional focus and motor learning: A review of 15 years. International Review of Sport and Exercise Psychology, 6(1), 77–104. https://doi.org/10.1080/1750984X.2012.723728
IPA note: example transcriptions were checked against Wiktionary entries and Wikipedia phonology articles that cite the Journal of the International Phonetic Association Illustrations for each language.