The Speech Genie Kickstarter story

Think in your new language from day one.
One new thing at a time. Never two.

In six months, you could be having real conversations in a language you don’t yet speak.
Not reciting. Not translating in your head. Not waiting until you feel ready. Actually talking — and understanding what comes back.
3,124 drawings, each made to carry a single meaning, native-speaker audio, and training that stays in the language you’re learning — sequencing designed over twenty years, so nothing arrives before you’re ready for it.
No translation · Nothing to memorise · No grammar drills · No scores · No streaks
| What it is | A language learning system designed around the way you learned your first language: meaning first, no translation, no grammar drills. |
| Already built | Evolved from a system in market since 2010 (10,000+ users); 3,124 drawings done; the core engine is running — you can play it right now. |
| You get it now | Something real lands the week the campaign closes — not in 2027. |
| Price | $199/year at launch; founding backers pay far less, 7 Oct – 4 Nov only. |
Thirty seconds of watching. Ninety seconds of doing.
That’s the watching. The doing is one click away ↓
You just watched someone else understand a language they don’t speak. Now it’s your turn — free, no sign-up, in your browser. Ninety seconds and you’ll feel the difference yourself.
Go and do that now. This page will still be here.
What you get if we hit our target. And when.
You are not only funding something. You are buying something — and part of it arrives within weeks of closing a successful campaign.
Every person alive has learned at least one language. You mastered yours to fluency before you could read, before anyone taught you a rule, before you knew what a verb was. The method you were born with worked. Then school handed you a different one.
Speech Genie recreates the method you were born with, rebuilt as software.
Your first ten minutes
What follows is a description of the first lesson, in the order it actually happens. It runs in your browser, and the demo in the “Play the Free Demo” link above is where it starts — within ninety seconds you’ll have a good feel for how it works. The whole exercise might take you about ten minutes.
You pick a language
English or Mandarin to start, depending on what you want to learn. More languages follow — and backers help decide what comes next.
You hear something, and you understand it
A voice says a phrase in the language you’re learning. Several drawings appear. You choose the one that matches.
No translation. No grammar explained. Nothing written down.
It starts as simply as “the door.” By the end of the course you’re following “Spread the butter on the chicken with the knife.” Still with nothing in between. Just meaning, arriving direct.

What Mandarin means here — and what it doesn’t
At launch, Mandarin means understanding and speaking. There is no formal teaching of the writing system — no stroke order, no radicals, no character lessons, and nothing you are ever required to read in order to learn.
But you won’t come out of it blind to the script. In FaceFonics® — where you’re watching how a syllable is physically made — the character appears beside the sound it belongs to, along with the pinyin: the Roman alphabet with tone marks, used to write Mandarin sounds. As you proceed through the course, pinyin and characters will start appearing in your Language-to-Body training as well.
Meet them enough times attached to a sound you already know, and they become recognisable.
Which is the same mechanism as everything else here. You aren’t taught the form. You meet it until it’s familiar.
The order is what makes that safe. An English speaker who reads pinyin first — before the sound is in their ear — sees a q, an x or a zh, says it the English way, and carries that mistake for years. Tone is worse: you cannot read a tone, you can only hear it and match it. So the sound always comes first, and the written form attaches to it afterwards. That way round, recognition costs you nothing.
So: ears first. You will understand Mandarin, and respond to it, before you can say it. You will speak it long before you can read it.
That is the order every Chinese child does it in, and the order that produces people who sound right.
Learning to write is a different animal — stroke order, radicals, thousands of forms — and it is not part of this campaign. Recognising what you are looking at is included. Writing it isn’t.
The Mandarin taught is standard neutral Putonghua — the Mandarin used in broadcasting and taught in schools, not one city’s local accent.
You hear a native speaker — then you hear yourself
You play the native speaker. You record yourself. You play both back.
And then you listen. Not close enough? Go again. Happy? Move on.
Nothing marks you. Nothing scores you. The comparison is the training — the gap between the two recordings is what your ear learns from, and your ear is the thing that has to change. Nobody ever learned a sound by looking at a picture of one.
It also means nobody has to watch you do it. Close the door and be terrible for as long as you need to be. Embarrassment stalls more adult learners than difficulty does, and it disappears entirely when there’s no one in the room.
Nothing is ever defined
Ask almost any course what a chair is and you get a word in response: 椅子, silla, chaise. Speech Genie shows you fifteen of them instead — armchairs, an office chair on castors, a folding chair, dining chairs — and lets you find the edges yourself.
Then it shows you a backless stool, and that one is wrong.
Which is the whole lesson in a single tap. A set where everything belongs teaches you membership. A set with one exclusion teaches you the edge — and the edge is what you actually need, because it is where one word needs to become another word.
That stool has one more thing to teach. It isn’t a near-miss at 椅子 — it’s its own word, 凳子, and the line between them runs along the backrest. English draws the same line in the same place: a stool isn’t a chair here either.
And neither language enforces it. Stand in a room with nothing but stools, say “pull up a chair,” and everyone knows exactly what you meant.
That is the part no definition reaches. A word list gives you one word and one gloss and tells you nothing about where the category stops — or how far it bends when the room makes the meaning obvious.
It’s how you learned the word the first time. Nobody defined chair for you. You met enough of them — and enough things that weren’t one — that the concept assembled itself. A translation hands you one word for one object. Range hands you the idea, edges included.
Nobody memorises their first language. Meeting a word enough times, in situations where its meaning is obvious, is how it gets stored — and that is a different activity from sitting down to learn a list.

Something we only learned by watching people fail
Our Chinese learners of English kept passing listening tests and then failing at listening.
They would score well, walk into a real conversation with an Australian colleague or a client in Delhi, and understand almost nothing. Test audio is slow, clear and spoken in one accent. The world is none of those things.
So the English course is built on international English — neutral American, mixed with British, Australian and New Zealand voices. That change on its own fixed a large part of the problem.
You did not fail at listening. You were trained on one voice, and then sent to a planet with thousands.
You start with the words that actually matter
Most courses drown you in vocabulary. Roughly a hundred words cover half of everyday conversation. Two thousand and you’re following 80% of daily life.
Speech Genie gives you the right ones, in the right order, and doesn’t waste your time on the rest.

You were never asked to do two new things at once
What you experienced in the short demo is the shape of every lesson, and the order is the whole point.
Understanding comes first, with nothing else asked of you: you hear, and you choose a picture. Then sound on its own — in FaceFonics you watch how a syllable is physically made, with no meaning to work out at the same time. Then producing: the native speaker and your own voice, played back against each other, with nothing left to decode.
Each step asks for one new thing.
That matters more than it sounds, because speaking a language is not one skill. It is at least five, running at the same time. Hearing where one word stops and the next begins. Recognising what they mean. Finding the words you want back. Putting them in an order the language allows. And making your mouth produce sounds it has never made before.
Every one of those is new when the language is new. Ask an adult for all five in the same moment and they freeze — and almost everyone who freezes decides they are bad at languages. They aren’t. They were asked for five new things at once, which nobody can do.
So the course never asks. Each skill gets built while the others are resting, and by the time they have to work together, none of them is new any more.
Twenty years went into working out that order. The drawings are the visible part. The sequence is the part that took the time.
This isn’t a plan. A lot is already built, and we already have prototypes for the new pieces.
Crowdfunding usually asks you to believe in something that doesn’t exist yet. This is a different proposition.
The method worked before there was any software
Chris wrote The Third Ear with no app, no drawings and no system behind it — just the principles, on paper.
Readers wrote back saying they had reached fluency in six months. In Spanish. In Portuguese. In Mandarin. People working entirely alone, from a book.
That is the part worth sitting with, because it means the method does not depend on us.
What it does depend on is having the right things to listen to, in the right order — and that is where most people working from a book run out of road. Real life will not hand it to you. Real life is random, and at the beginning almost everything you hear means nothing, so immersion mostly means standing in noise you cannot use yet.
That is what twenty years of work went into. Not to do the work for you — you still do the work — but to make sure every minute of it lands.
Then it was built into software
Kungfu English has been live on iOS since 2010. Chinese speakers use it to teach themselves English — on their own, at their own pace, with coaching only if they want it.
It was the number one top-grossing app in China’s App Store when it launched, and more than 10,000 people have paid for it since. Not trial sign-ups. People who paid.
One of those learners is on this team. Meihan Chen learned her English from Kungfu English while it was being built, and has run its marketing ever since. When we say the system teaches people to teach themselves, she is an example of what we mean.
From 5% passing to 67% passing — in just two months. One change of method, thirteen times the result.
Then a classroom showed us what was missing
In 2015 Chris designed the course and led the teaching team at Bainian Vocational School (百年职校) in Beijing. The programme ran on two legs. In class, traditional TPR — a teacher demonstrating with her own body, students responding to what they heard — which built the physical connection to the language, and the motivation to keep going. At night, the students went into Kungfu English on their own: one to two hours, most nights.
We say what produced that, because it matters: a teacher in a room and software, together. Not the software alone. One cohort, one school, with Chris leading the team on the ground. It isn’t a controlled trial. It’s what happened when the method met a real classroom.

And it left us with a problem, which is the reason this campaign exists.
The classroom was doing something the software could not. TPR has two halves: someone demonstrates, and you respond with your body. At Bainian the teacher carried both. The next eight years went into making the software carry the first half, and that meant three things at once: designing and coding it, building the curriculum, and creating the materials it teaches through. All of it has been in the market since 2018.
This campaign is about the second half.
Speech Genie is the fourth generation of the method
2006 — the method set out in The Third Ear. 2010 — built into software. Kungfu English goes live on iOS. 2018 — the line drawings go in. 2026 — the scenes start to move.
Same ideas. Same discipline. Now being rebuilt so it works in any language, for anyone.
3,124 drawings. Nothing written. Any language.
We commissioned a library of line drawings that teach meaning directly — no writing, no translation, no culture.
Commissioning them took six months. Working out how to teach a language through pictures and software — which images, in what sequence, doing what — took twenty years.
Anyone can commission illustrations. Knowing exactly which 3,124 situations will carry a person from zero to following “spread the butter on the chicken with the knife” — with nothing in their own language to hold onto — is the actual asset. The drawings are what that knowledge looks like once it’s been drawn.
Most of them aren’t pictures of things. They’re pictures of situations — mini-stories, a whole event held in a single frame. That distinction is why the system works at all. A flashcard can show you a chair. It can’t show you the boys pushing their bicycles on the narrow path.

Between them they carry everything words normally have to carry: objects, actions, motion, direction, sequence, comparison, cause, even emotional state. A dashed line circling a house is around. A figure mid-stride with an arrow is walking to.

Because nothing is written on the drawings, the library transfers to any language on earth without being redrawn. It’s the hardest-won thing we own, and it’s finished.
And the system barely needs to know what language you already speak
Every other course needs to know exactly. They’re built as pairs — English for Spanish speakers, English for Japanese speakers — because the teaching runs through your mother tongue: the instructions, the explanations, the grammar, all of it.
Speech Genie uses a little of your own language to get you started, and then stops. There is no grammar translation anywhere in it. Nothing is explained in your language, because the drawings do the explaining — and a drawing means the same thing to everyone looking at it.
One course per language. Not one per pair.
That isn’t a convenience. It’s the whole reason a library with nothing written on it had to exist — and the reason working out its contents took two decades.
Which is why the evidence above is evidence for you
Almost everything we’ve proved so far, we proved in one direction: Chinese speakers learning English.
That would be weak evidence if Speech Genie taught through your mother tongue — a course built for one language pair tells you nothing about any other.
But nothing here runs through your own language. So nothing in those results depends on the pairing. What was being tested over the last decade was the method itself.
And the readers of The Third Ear went the other way entirely — English speakers into Spanish, Portuguese and Mandarin, with no software at all. Same ideas, different delivery mechanisms, results every time.
And if English is the language you can read but not speak, this was built for you first.
You are reading this page, which is the problem stated in one sentence. Years of study have given you enough English to follow an argument written using English, and not enough to say what you think out loud, at speed, to somebody waiting for an answer. Nothing is missing from what you know. What is missing is the road from knowing to saying — and the method you were taught never built one, because everything arrived through your own language, and it still has to travel back through your own language before it reaches your mouth.
That is the gap the Bainian result was measured across. Those students were not absolute beginners. They were exactly where you are. And it is why the English course is the one thing here that doesn’t wait until 2027 — it is a re-sequencing of material that our learners have been using since 2010, and it is in your browser the week the campaign closes.
And the core game engine is running
The core interface that carries all of this into a cross-platform game is built and working. Not a mockup — running software, some of which you just played.
There is more engine to build, and the UI needs a rethink. That’s part of what you’d be funding.

If you want to see how the whole thing fits together — speechgenie.co/full-demo →
Everyone around him was failing. He wanted to know why.
Chris Lonsdale learned Mandarin from scratch at 22, in the early 1980s, and reached native level. Then Cantonese, fluent in under six months.
That should have been the end of it. Instead it started a thirty-year question, because everyone around him was doing the same work and getting nowhere. Same effort. Same hours. Almost no results.
The difference wasn’t talent. It was method.
The one almost everyone is handed was designed in Prussia in the 1780s, to teach Latin and Ancient Greek — languages nobody was expected to speak. It suited administrators marking papers. It has never suited anyone trying to talk. Two and a half centuries later, most courses still run on it.
Well over one billion people are learning a language right now. School-based failure rates run 70 to 80 per cent. Even if only one in three falls short, that’s more than 500 million people a year who don’t get where they were going.
And to be clear about the language-learning apps
The category works. Duolingo has 140 million monthly users and nearly 13 million people paying for it. Nobody needs persuading that people want this.
What it hasn’t produced is speakers. Duolingo’s own founder has been straight about that:
“I won’t say that with Duolingo, you can start from zero and make your English as good as mine. That’s not true. But that’s also not true with learning a language in a university, that’s not true with buying books, that’s not true with any other app.”
— Luis von Ahn, TechCrunch, May 2021
He is right on both counts, and the second is the more interesting point. He is not describing a gap in Duolingo. He is describing a gap in the whole category, and it has been there long enough that people stopped noticing it.
None of which is a criticism, and we’re not going to pretend otherwise: getting hundreds of millions of people to start, and to come back tomorrow, is a genuinely hard problem and the category solved it.
Real interaction while thinking in your new language is the part nobody has taken. That is what Speech Genie is specifically designed for.
If you have ever had your mind go blank when you try to say something in another language, you’ll know that thinking in that language is the way out.
And no, you’re not too old
Somebody told you children are sponges. That their brains are built for this and yours stopped being built for it a long time ago. Maybe you tried at school, it didn’t take, and you quietly filed yourself under not a languages person.
Here is what is actually true about children.
A child takes about five years to hold a decent conversation. Five years — surrounded by the language every waking hour, with no job, no bills, and no idea they are supposed to feel silly. They are not fast. They are relentless, and nobody laughs at them.
You have things a four-year-old does not. You already know what a sentence is for. You know a thousand things about the world that a child has to learn alongside the words. You can concentrate deliberately.
What you don’t have is five free years. What you do have is the fear of sounding stupid.
Those are the real obstacles. Not your age.
The largest study ever done on this put the same test in front of two thirds of a million people. It did find a decline — but read carefully what declines. What drops away in late adolescence is the rate at which you take on new grammar. Not the ceiling. The authors are explicit that those are two different questions, and that the answer to one doesn’t give you the answer to the other.
So we went to their released data and counted people, instead of reading curves.
Of the 9,720 people in it who started English at twenty or older, 3,096 scored above 90%. Nearly one in three. The share thins as the starting age rises and it never reaches zero — people who began in their fifties are still clearing that bar.
And 16,474 native English speakers scored below it.
Those two groups overlap, and they overlap heavily. Thousands of people who started as adults scored higher than thousands of people who grew up speaking English. So the age at which a person starts learning tells you a lot less about where you end up than what you have been led to believe.
Nor is immersion the explanation. Of those 3,096, more than a third had no recorded time in an English-speaking country at all — and the older they started, the more likely that was. Past fifty, most of them had never lived there.
One more line from the same work that almost nobody quotes: native speakers are still getting better at their own language until around thirty. Nobody is finished at seventeen. Not even the natives.
Which means, if you want to learn a new language you can succeed — but you do have to go about it the right way.

Hartshorne, Tenenbaum & Pinker (2018). The counts are ours, taken from the dataset the authors released; the definitions behind them are set out in full at The full argument →. The same lab’s 2021 follow-up, across 1.1 million people, places the decline between 17 and 19 and notes that their original model was biased towards finding a sharp one.
Thirty-seven million people have watched Chris explain the method
The TEDx talk, “How to Learn Any Language in Six Months,” is the most-watched TEDx talk ever given on language learning — 37 million views on YouTube alone.
Then a book, The Third Ear. Then a second, written in Chinese: a bestseller across mainland China, Korea, Taiwan, Macau and Hong Kong.
If you’re one of the people who watched that talk and thought someone should build this — this is that thing. Built.
And if you’ve never seen the talk, click through and watch it now. We’ll be right here when you come back.
Chris knew what he wanted to build in 2010. The technology didn’t exist.
He wanted a patient, endlessly available language parent. Someone — something — you could practise with at two in the morning, with nobody watching and nothing to be embarrassed about.
The missing piece was software that could work out what a learner actually meant.
John Ball spent thirty years on exactly that
John, a computer scientist and cognitive scientist, worked on a single question: what would it take for a machine to handle language the way a human brain does?
After a collaboration with Marvin Minsky he reached his answer. Meaning isn’t assembled from rules, and it isn’t guessed from frequency. It’s held as patterns. He built that into what he calls Patom theory, then grounded the theory in Role and Reference Grammar — the linguistic framework that maps how the world’s languages carry meaning — so it could be written as working code. That code now exists.
Not scored. Understood.
Your Genie is being designed to observe your progress and guide you — then, when you’re ready, understand you.
Ultimately, that’s what you’ll be talking to: a meaning engine, built on thirty years of work on how humans hold meaning, and tuned to the way your brain already does language.
Not marking your accuracy. Understanding you.
Four things this project builds
More language — double the vocabulary, and the scenarios and drawings that carry it. Scenes that move, so you stop pointing at a picture and start acting inside one. An interface that teaches without words, because a course with no translation has nothing else to explain itself with. And it has to run on whatever device you already own, or none of the rest of it reaches the people it was built for.
They are in build order. The first is the most work. The second is the most visible. The third already works, and this campaign is what takes it to what it could be. And the fourth decides how many people ever get to use any of it.
1 · Double the language
The library covers roughly 1,000 of the highest-frequency words in English and Mandarin — the ones carrying most of what anyone says in a day. All 3,124 drawings built on them come across into the new game.
This campaign then adds 1,000 more: colours, shapes, numbers, directions, and many more objects and actions. They arrive the way everything here does — not as word lists, but built into new scenarios and newly commissioned drawings.
That takes Speech Genie to roughly 2,000 high-frequency words — enough to follow language for around 80% of daily life.
We aren’t inventing a capability here. The hard part isn’t the drawing. It’s knowing which situations to draw, in what order, and what the software does with them — and that is the twenty years, already spent. Commissioning the drawings themselves is a known, bounded job: the original library was produced in six months. And the course-design software that assembles them into lessons is already migrated from Kungfu English and running.
2 · The scenes come alive — and you act in them
This is the part we’re most excited about.
Today the drawings are still. You hear a phrase and find its meaning in the picture. What the line drawings had to prove was something narrower than Bainian: that the step still works when the teacher’s body is replaced by a picture. Eight years in the Chinese market is the answer to that.
But understanding is only half of language.
With this funding the scenes will also become performance stages. The things inside them — the cup, the door, the chair, the person — become things you move. You hear “put the cup on the table” and you don’t point at a picture of it. You make it happen.
Start talking with your fingers so your mouth can catch up.
Every child does this before they say a word. They respond with their body long before they can produce speech, and the responding is what wires the language in. It’s the step adult courses skip entirely — and it’s the one that makes everything afterwards easier, because this is where you start thinking in the new language instead of translating into it.
None of which is new. Total Physical Response has been researched and taught since the 1960s, and it works. But it has always needed a room, a teacher standing in front of you, and a group moving in time with each other.
The drawings already do this. And now we are automating it. One learner, their own pace, nobody watching, no classroom to get to — until they’re ready for a live language parent, and can hold their own with one.
And the scene answers you. Get it wrong and nothing is marked. Your Genie speaks again — I said this; you did that; try this — in the language you’re learning, and you understand it, because you have just watched the difference happen in front of you. Being wrong stops being a dead end and becomes the next thing you understand.
We have a working prototype and it’s promising.


3 · You stop learning from your own language and start learning in the new one
This is the largest piece of work in the campaign, and the least visible.
Kungfu English always knew what language its learners already had. Whenever something needed explaining, it could say it in Chinese. Speech Genie can’t — it has to teach someone whose first language it doesn’t know. Past a little guidance at the start, nothing can be explained. It can only be shown, in an order that builds.
Which puts the whole weight on the interface. Where a thing sits on the screen, what moves and what waits, what invites a tap — in a course with no translation to fall back on, those aren’t styling decisions. They’re teaching decisions. The interface is the method.
We are still making those decisions. What you played in the demo is real, working software, and it is not the finished design.
4 · It runs on whatever device you already have
Kungfu English is one iOS app. That was enough for one market. It is no use at all to most of the people this is for.
The learners who need this most are not the ones holding the newest phone. They are in classrooms and on buses in places where the device is whatever was affordable, and where the browser is the only app that matters. So Speech Genie is being built to run across platforms, and to run in a browser, on hardware that is already years old.
This is the least glamorous of the four, and it decides who is able to use the other three. A method that only reaches people who can afford the right phone has not solved the problem it set out to solve.
We have prototype builds on two paths — Unity, and JavaScript/HTML. Which one carries the product is not decided yet, and making that decision well is part of what this campaign funds.
And you are in it before we are sure
This one is not something we just build on our own and deliver to you. It is something we cannot do without you.
The drawings are the settled part — a decade in one market has told us what they teach. What isn’t settled is everything around them.
A screen has to tell you what to do without using your language to do it: where to drag, what counts as an answer, what just happened when you got it wrong. Every one of those is normally a sentence, and we don’t get sentences.
Then there are the words that aren’t things. A chair can be drawn. Around, before, instead, almost cannot — not directly. They get built out of arrangement, motion and sequence, and whether a particular arrangement reads as before rather than behind is not something we can settle by agreeing among ourselves.
And there is what language is doing rather than what it says. Give me the cup and could you pass the cup point at the same picture. The difference between them is what makes you sound like a person instead of a phrasebook, and it is the hardest thing to teach when nothing can be explained.
We can’t test any of it from the inside. We wrote the conventions, so we already know a dashed line means around — and you can’t un-know that in order to check whether it teaches itself.
So backers get the Alpha, and we ask specific questions. Did that screen tell you what to do, without resorting to words? Where did you hesitate? When two things looked almost the same, which did you pick, and why?
Not feature requests, and not a vote. It’s the one measurement we can’t take from the inside, and we’d rather use it to shape the interface and the order things are taught in than hand you something we settled on by ourselves.
What it costs, and why we can tell you that with a straight face
Speech Genie will be $199 a year.
That number needs context, so here it is. The system underneath it — Kungfu English — sells in China for RMB 6,980. About US$980. On sale since 2010, more than 10,000 people who paid it, and a profitable business at the end. We are not guessing what this is worth. We have been selling it for a decade and a half.
The other comparison is the one you would actually make. Six months with a tutor, three sessions a week, costs $1,400 to $2,900. A classroom course costs more and runs on someone else’s timetable.
Speech Genie is $199 for the year — at six in the morning, on the train, and as many times as you need the same sentence without anyone sighing.
Backers pay considerably less than that, and every founding price on this page exists only while the campaign is open — from 7 October to 4 November 2026. They will not come back.
And there is a version where you never pay again
During this campaign only, you can take your language for life — including every capability we ever add to Speech Genie itself. Not a year, not a discount on a year. Once, and then never again.
To be exact about what that covers, because a promise this open needs an edge: it is the Speech Genie course in the language you chose, for as long as the product exists, with every improvement we make to it. It is not a claim on other languages, other courses, or products we have not built yet.
It is capped, and it ends when this campaign does — 4 November 2026. After that the lifetime tiers close and do not reopen. If you already know you’re going to do this, it is the tier to take.
You don’t wait until 2027
Whatever you’re learning, something real lands when the campaign closes.
(Kickstarter charges cards at the funding deadline rather than the moment a goal is reached, so confirmation comes first and delivery follows on close.)
Learning Mandarin — your Rhythmic course. Three volumes, 950-plus words and phrases, native-speaker recordings, printable phrase books. A finished product we have been selling for years, yours to keep, with no expiry.
It is not Speech Genie. It’s a different course, built on Georgi Lozanov’s Suggestopedia — rhythm and music used to carry language in below the level of conscious effort. Different mechanism, same conviction: the learning that lasts is the learning you didn’t have to force.
Learning English — the core English course itself. Not a substitute: the actual Speech Genie English course, playable in your browser. It is a re-sequencing of material that our learners have been using since 2010, so it does not wait on the rest of the build.
It will be rough. The interface won’t be finished, there are no apps yet, and you’ll find edges we haven’t smoothed. That is the point — we said above we would rather build this with backers testing it than hand you something we settled on alone, and this is that promise starting early rather than a compromise.
Neither costs you a day of your subscription. The clock starts when Speech Genie ships in 2027 — everything before that is on us.
A link lands in your inbox as soon as Kickstarter hands us the backer list, usually within a week of the campaign closing. One click and you’re in.
Worth doing now: whitelist speechgenie.co so it doesn’t land in spam.
Which also means the goal is not just our milestone. It is the moment everybody who has already backed gets something real.
Stretch goals
Stretch 1 · Real life, and another thousand words
Role plays for the situations you’ll actually be in: ordering in a restaurant, dealing with a bank, asking directions, handling the moment when someone speaks too fast.
And because that’s how we build vocabulary — never as word lists, always inside situations — these scenarios carry another 1,000 high-frequency words with them. That takes Speech Genie to roughly 3,000 words: the point where you’re following almost everything you’ll hear in an ordinary day.
Same drawings, same method. No new system to learn — just the language you already have, put to work in the places you’ll need it.
Stretch 2 · The roles flip
In the animated scenes you already answer back — with your hands. Past our stretch goal you answer with your voice. You give the command.
You say “kick the ball over the table” — and your avatar kicks the ball over the table.
Not a tick. Not a score out of ten. The ball moves, because of something you said in a language you’re still learning.
And your Genie is delighted — because that’s what people do when they understand you. Not points. Not a streak. The ordinary human pleasure of getting through to someone, which is the thing you were chasing all along.
Worth noticing what that takes. To act on that sentence the system has to work out who is acting, what the action is, which object is affected, what spatial relation is being asked for, and which object that relation refers to. Get any one of them wrong and the wrong thing visibly happens.
That is understanding, in the only form that can be checked. The ball goes over the table, or it doesn’t.
Role reversal starts in the core build, where you answer with your hands. What the stretch goal buys is the voice — built by integrating speech recognition with John’s meaning engine, and the first step of something larger still: two-way conversation. We’re naming the first step because that’s the one we’d be committing to.
Who’s building it
The people behind it.
Product & Methodology

Chris Lonsdale
Psychologist, linguist and educator. Creator of “Learn Any Language in Six Months.” TEDx speaker, 37M+ views. Built Kungfu English.
Co-founder · Language technology

John Ball
Cognitive scientist. Thirty years on how meaning works, including a collaboration with Marvin Minsky. Creator of Patom Theory.
Strategy & Marketing

Meihan Chen
Marketing lead for Kungfu English for 15+ years, and one of its first learners — she learned her English from the system while it was being built. Deep China consumer-education experience. Grew Kungfu English to more than 10,000 paying users.
Community & Partnerships

Beth Carey
Former senior manager at IBM and Fujitsu. CEO of Pat Inc. Founder of Women in AI (Australia).
Co-founder · Japan

Bryan Keniry
20+ years in Japan, fluent in Japanese — which he learned, and then taught, using Total Physical Response: the method Language-to-Body automates. Creative software design and problem-solving in complex systems. Ex-IBM. The animated Language-to-Body core concept, and the initial prototype, are his.
Ninety seconds, no sign-up. Feel it for yourself.
The whole argument on this page comes down to one thing you can test in your browser right now.

