Rendered at 07:23:06 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
skolos 11 hours ago [-]
Interesting that this is here. I used whistle (and bunch of other things) to take ownership of my echo show. It now doesn't dial to Amazon at all - it does all processing locally with its own CPU and connects to my homeassistant for home automation. My initial setup involved qwen asr (1.7b model) running on rtx 5080. Compared to that, whistle was really bad (out of 170 messages, qwen recognized correctly 168, whistle - 70), but I adjusted whistle to work like jev - instead of free form transcription it recognizes only select set of templates (I trained tiny network with 10,000 generated utterances to translate whistle final state to probabilities within templates). The precision went up to 164/170 - almost matching qwen. By the way - I'm speaking with heavy accent.
mooman219 9 hours ago [-]
This feels like we're going back to the past again with CMU's Sphinx4 in Java. It worked way better than I would ever have expected it to for being more than a decade old. It relied on the user defining a grammar of valid words and different flows through a standard format (Java Speech API Grammar Format). I wonder if we'll approach that again for these models just like how MCPs act like WSDLs in spirit. Great results getting whistle working so well for you!
#JSGF V1.0;
grammar commands;
public <action> = open | close | delete ;
public <object> = file | window | application ;
public <command> = <action> [the] <object>;
skolos 9 hours ago [-]
Just ran some tests on my phrases:
- my whistle setup (tested can run on echo show with <1s response): 206 correct out of 208
These were not tested on echo show yet - on my pc for now:
- Vosk small, phrase grammar: 184/208
- Speech-to-Phrase 1.4.3 (Kaldi): 118/208
- suggested sphinx: 63/208 and on some examples took 13s on i7 14700k pc
looks like my customized whistle works better for me than these alternatives. But more testing wouldn't hurt.
ianjbutler 7 hours ago [-]
Not op but I think maybe the connection was more grammar vs your “select set of templates”.. observation being that structure beats no structure regardless of the base tech, and that completely structure-free interactions are generally not what we want anyway
Multicomp 6 hours ago [-]
WSDL - now that's a name I've not heard in a very long time!
XML Web Services were awesome, but the SOAP and XML (where the hard parts that should have been given some batteries included defaults) I think was too much of a boat anchor to overcome.
And then the ruby/rails wave came and made 'rest' JSON the Silicon Valley hearthrob, and WSDLs were out like yesterdays trash.
Are they still used in businesses IRL? Yes! I worked for one that has probably (probably) moved off of them by now, but as recently as 2023 there were still some bank-facing power with SOAP RPC .asmx endpoint servers.
I follow hypermedia.systems and htmx/datastar/alpine.js because I'm still trying to get back that powerful _web_ tech advancement rather than squeezing everything into the javascript client side workaround. Typescript is great, a good poor man's F# and leagues ahead of ES3 (which I dabbled in / torture LLMs with so my old retrocomputers can still do web things), but its still bringing along so much baggage that could be lighter weight for older devices AND fits the tech utopianist promise of the original web (which is half of why I play with computer stuff).
miki123211 2 hours ago [-]
Poland runs on WSDL, SOAP, XML Schema, XSLT, XML Forms, all of it.
By law, every single government form you're able to file is supposed to have an XML Schema available in a centralized registry. This regulation is widely ignored in practice, particularly by local / municipal governments.
This used to be more important in the past, as such a form could automatically / semi-automatically get an entry in the EPUAP form catalog and be made electronically fillable (EPUAP being the now-deprecated centralized government bureaucracy portal basically). As far as I understand, the way that worked was through XSLT and XML Forms. There was some weirdness about each document having a fillable / form view and a preview, I think XML Forms was somehow used to generate the XHTML form, while the XSLT sheet could only generate an XHTML preview of a complete, filled-in version. There were also some custom annotations for auto-fill and such, that was partly done by most documents relying on standardized schemas for entities like "person", "address" or "company".
Since we moved to E-Deliveries and lost a central place for these forms to live in, this is (AFAIK) a bit less important and less common, but internally, things are still XML. If you're filing something like an ID renewal application, even through a newer, more user-friendly frontend, it's still just XML underneath. If you sign something through podpis.gov.pl (the standardized e-signature solution for government paperwork), you can even download that underlying document, both signed and unsigned, and see what that XML is. The pre-signing document preview generated by that site still comes from the XSLT I believe. The signatures themselves are, unsurprisingly, also done via the XML signatures spec.
Incidentally, the European E-Delivery system itself also relies on XML, WSDL and Soap pretty heavily. For those unfamiliar and/or not in Europe, it's basically "email but for the government", with all the guarantees and legal obligations of physical mail, cryptographically-attested confirmation of receipt, proper identity verification and assurance, cross-provider address portability, deployed to a lesser or greater extend in many EU countries and set to replace physical mail.
cco 2 hours ago [-]
I found parakeet to be very good and pretty darn fast, ~20-50ms, accuracy around 9/10 whenever I've run tests. In practice the error rate is not too big of an issue because the output is fed into GLM 5.3 so it'll figure it out if it sees "heather" instead of "weather".
Much larger though, I think I went full precision and its around 2GB.
stronglikedan 10 hours ago [-]
> By the way - I'm speaking with heavy accent.
I chuckled at this because my inner voice had an accent as I was reading your comment, due to your writing style.
yuchi 10 hours ago [-]
Sorry, curious non-native speaker here. Which accent? And which telling patterns made you think of it?
Nition 10 hours ago [-]
Not the person you're responding to, but as soon as I read stuff like "I trained tiny network" I tend to imagine Russian. It's the missing "a". Same thing happens in "I'm speaking with heavy accent".
meredithbloom 2 hours ago [-]
I'm not a native speaker but my inner voice also switched to Russian accent very quickly!
Zacharias030 10 hours ago [-]
Missing articles in English is the usual give away for someone with a slavic native language such as Russian.
ASalazarMX 10 hours ago [-]
Excuse me, could you write slower? I couldn't understand you.
catlifeonmars 5 hours ago [-]
How did you load your own programs on the show? I own an unreasonable number of old echo devices (a couple shows too) that are sitting in a box. This sounds like the perfect use for them!
psyclobe 4 hours ago [-]
Huh I tried that was slow as all hell to respond
schappim 10 hours ago [-]
Have you done a blog or YouTube about this ?
skolos 10 hours ago [-]
It is all custom made and not very reproducible. I'm working on reproducible setup and once it is done will publish it here: https://blog.kvit.app
mrguyorama 10 hours ago [-]
If you are purposely limiting yourself to select templates, even fairly complicated templates, then classic voice recognition is perfectly sufficient.
With a restricted grammar, built in Windows voice recognition, all on device, has managed this exact use case quite well for over a decade. I used it to try and build a clone of the various paid apps that allow you to issue orders to Arma soldiers with voice commands
INTPenis 14 hours ago [-]
I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.
I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.
But every single sound he makes with his mouth ends up on the page too.
ComputerGuru 14 hours ago [-]
Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.
For essentially infinite and fast dictation I use https://github.com/cjpais/Handy on Parakeet streaming (cohere is far better, but slower and has a token output limit so you cant ramble for many minutes). And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words. I just weave instructions into my writing. I understand this requires technical know-how, but for those with it, this is the best solution I have found to long form writing without my hands.
lnenad 10 hours ago [-]
I literally today pushed v0.1 of my dictation app that does both in a single package. You can connect either to a cheap LLM but I've set it up out of the box to use Qwen3.5 9b which does the job well for free, and everything stays local, no telemetry. You can also use Claude/OpenAI/Local providers. https://github.com/lnenad/lipwise
drzhouq 9 hours ago [-]
Difference with voice ink?
hatsix 4 hours ago [-]
Open source, multi platform
mkbkn 5 hours ago [-]
Does the AI processing work fast or consume less resources on 4-5 years old laptops with 8GB RAM?
lnenad 14 minutes ago [-]
For slower machines it's much better to use a cloud model, but Qwen3.5 4B can do the job and is fast enough on older machines.
johanvts 10 hours ago [-]
Neat, thanks!
dv35z 12 hours ago [-]
Another plug for Handy, and wanted to share something cool about it.
You can set it "Push to talk" mode (like a walkie-talkie radio), and when you're done talking and release the button, it can paste the text into any text field.
You can even replicate ChatGPT voice conversation mode, by having Handy as your speech input, and then (I forgot the extension) enabling a speech-to-text model for OpenCode. Surprisingly relaxing flow for certain tasks, like tweaking a website's styles.
mongrelion 2 hours ago [-]
> And then just do a cleanup pass with a cheap LLM
I am especially interested in this part. Could you please share the prompt you are using to instruct the LLM to clean up the dictation? Thanks in advance!
flockonus 10 hours ago [-]
Agreed, record and transcribe however you can and then use a LLM to clear out.
I prompt it to:
"Attached (or underneath) is the transcript of a self recording i've done with tons of rambling and some incorrect words transcriptions, please do a pass clearing out and arranging any typos or possible misunderstandings. Keep original in parenthesis when not sure if it's a misunderstanding. Do not summarize or alter the nature of the content, simply tidy the transcript."
Barbing 9 hours ago [-]
Original in parentheses is very interesting.
apitman 10 hours ago [-]
Love Handy. It has become a core piece of how I use computers.
In case it's helpful to anyone else using it, at first it felt a bit slow to me, because there was a noticeable pause after I finished a message before it would quickly type it all out. I changed the input method from direct to clipboard and it's way faster now, almost instantaneous.
Technical knowhow? Handy is a gui right? (Also Parakeet is great I use it everyday)
chr15m 5 hours ago [-]
Thanks for sharing Handy!
Imustaskforhelp 12 hours ago [-]
+1 for handy and then using LLM's for the cleanup pass, though what are your observations on feeling as if sharing that output though?
Because I have seemingly mixed opinions on it, on one hand, I did put the effort but on the other, the output is AI generated so I am unsure about sharing it with others (because they might think its AI generated)
Do you use it for very small edits (removing just the uhhm's?) or for slightly more edits.
The way that I use it sometimes is that while thinking, I will write something which can sometimes make me feel as if a better re-write can better explain my thoughts or rephrasing it as such. For example. I will think about X topic, connect it to Y, then try to add some more points about X again.
I found LLM's to do a really decent job at generating the final outputs as such, but as I said, I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated and the end user not knowing if I put an actual effort into creation of it or not.
Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it.
asa123 12 hours ago [-]
what do you mean:
"what are your observations on feeling as if sharing that output though?"
and "Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it."
could you rephrase the question
for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.
Imustaskforhelp 4 hours ago [-]
> for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.
Yes, for doing extremely mild edits (like just removing uhh's etc.) this might be true
but I sometimes feel as if writing can allow me to shift paragraphs, so if I am writing para 1, para 2, I can shift back to para 1 and write another sentence in it and edit some parts of para 1 to include that point
but when I am doing STT, although I can move towards the other para, I find myself just speaking in a complete flow and just write in para 2 only.
Thus when I ask AI to write, I would prefer it to move the statements to appropriate paragraphs and in just general, create a more comprehensible viewpoint from all the STT text that I had written.
This does generate AI generated text which can be detected as such. Uploading it on blogs makes me feel as if people might read what they might consider "AI slop" and so the ethics part (as I myself don't wish to read AI slop)
the problem with AI written or edited texts is that I am unsure of how much effort the other person has put in (just a single prompt or a detailed thought was put in), and I feel as if, others feel the same way.
Should one try to show the rough draft as well to try to show that it was an effort which was human generated or that human effort was used, but that means having a proper disclosure that it was AI-generated/AI-assisted, which I feel as if offputs a lot people (including me) as because of the above logic, that there's still friction for the user within testing if real effort was put into place and I am unsure how effective sharing drafts of it could be.
I don't want my blogs to be tainted and treated as AI-slop because I care about them so I am unsure of what to do. I have multiple things that I have written which if I pass through AI can create some meaningful blog piece but as it stands, they are rough drafts and I find myself putting low efforts or being lazy in actually editing them myself as well (and potentially putting in multiple hours) when AI can be used to help create a more polished version just as well and get across my point.
nonmaskable 11 hours ago [-]
[flagged]
sy135673 2 hours ago [-]
[flagged]
nvtop 13 hours ago [-]
I've been using Gemini Desktop App purely for dictation. It's a miracle! For the first time in my life I'm blown away by the quality of my (heavy accent) speech recognition. Just be sure to disable the "speak to window -> reasoning" option to make it purely dictation and stop from writing whole emails for you.
islewis 11 hours ago [-]
> I don't think the challenge with speech to text was size of the binary
The usecase for small models like this is making on-device STT/TTS more accessable. This is important if your usecase is sensitive to either privacy or latency, but this comes at the cost of quality.
My experience has been that these small TTS models are unexpectedly good if your audio is in distribution (western accents, higher quality audio, common vocabulary), but pretty quickly degrade as you move outside of that. They often dont support more complex features such as diarization, multilingual, or realtime streaming either.
bob1029 6 hours ago [-]
> I don't think the challenge with speech to text was size of the binary
It's a very compelling aspect of the problem.
If you can get a model under certain size thresholds, that means you can eliminate latency domains. For example, if the model is able to fit entirely inside L3, the latency of servicing requests drops by an order of magnitude (or better) compared with a model that resides in L3+DRAM.
Retr0id 6 hours ago [-]
It's true that L3 is an order of magnitude lower latency than DRAM, but that's mostly going to translate into a throughput difference rather than a latency difference, when looking at the system as a whole.
boplicity 14 hours ago [-]
There are many different challenges, each requiring their own solution. I, for one, really miss the old Google Assistant on my Android phone. It would very reliably play most songs that I wanted to hear on Spotify. Gemini fails at this almost every time, and is significantly slower. It's actually a difficult problem, as the songs people want to hear are regularly being released, are often associated with uncommon names, or have words in unusual orders, so normal LLM style tools just don't cut it.
xp84 12 hours ago [-]
This reminds me of how the thing iPhones had pre-Siri (so we're talking pre-2010), which was entirely offline, did a better job than even the most modern thing at "Play [one of the finite set of songs in my library]." I sometimes get absurd matches from bands I've never heard of, when the right answer is something right there in my library.
raddan 12 hours ago [-]
It's odd that the matching algorithm does not simply prioritize the music already in your library, but I see something similar in other domains where machine learning/information retrieval is used. E.g., in Apple Maps, I might have the map centered over my location and type in a restaurant nearby. Often, Apple Maps will find a restaurant with the same name on the other side of the country. This strikes me as an easy thing to fix (and Apple Maps has had this bug from the beginning), but if somebody knows something abou this, I'd love to know. Maybe it's harder than I imagine.
xp84 8 hours ago [-]
I know it makes me sound like a lunatic to anyone listening, but I always use the most condescending monotone with Siri in the car, like this. "Directions to Wal-Mart,... in... Beaverton,... Oregon." Doesn't matter if I've been to that Walmart 197 times since Apple Maps 1.0 came out, because if you don't be as explicit as , 1 out of 10 times, it'll decide that you must mean a random Walmart 12 states away. Or like "Wall Plastering Incorporated."
And even when it's getting it right, and if there's only one Walmart in Beaverton, it still to this day needs to ask "One option is Walmart on Expressway Road in Beaverton...." Maybe it's correct in its 0% confidence level there, since it's so bad, but... I don't get how you could design something that bad, even before LLMs existed. I feel like I could do better, even using their Speech-to-text engine, with the processing backend built of pure regexes and if/elses.
rolosa 9 hours ago [-]
It's just useless for me now, the change happened some 5 or so years ago.
"Hey Siri, play [song]"
Leads to, take your pick:
- "You'll need to unlock your iPhone first."
- "I couldn't find [song] on Podcasts" (??????)
- "Playing [a totally different song]"
- "I couldn't find any music by [song, but it thinks it's a band]"
- "Playing music by [song, again it thinks it's a band]"
xp84 7 hours ago [-]
Yes!! I get all of those. My favorite is that 10% of the time it asks me on what app I want to play the music (Oh, Apple, you're suddenly deeply respectful of competing on an equal playing field?).
And don't forget whatever the current phrasing is for "I'm sorry, my shit's all fucked up" and "My network connectivity had a blip and I'm unwilling to retry" and "Even though I have on-device STT models, and now LLMs too, and an on-device database of your music, which is downloaded, I won't bother without the cloud.
NamlchakKhandro 4 hours ago [-]
so what's your problem? sell your apple products. I Know... you're from california, so this is heresy.
but once you calm down and stop hyperventilating from my suggestion, you'll see that the only reasonable and pragmatic course of action is to move to android and linux.
hbn 12 hours ago [-]
Settings -> Accessibility -> Side Button -> under "Press and Hold to Speak" choose "Classic Voice Control"
mrguyorama 9 hours ago [-]
This is because, in the case of a restricted set of possibilities, voice recognition circa 2000 was actually very very good.
If you can do something with an extremely limited vocab, voice recognition was fine using off the shelf microchips in the 70s, where you wired in a microphone connection and had discrete pins for output actions.
LLMs are basically only useful for utterly free form transcription, but that doesn't actually help you turn that into tasks to perform and parameters for those tasks
The core "problem" in voice recognition is that freeform speech is an abysmal UX paradigm and provides zero discoverability, and LLMs IMO have not improved the situation of actually doing anything with the resulting text.
The other day I tried to prompt Gemini 3 times to tell me what the heck the business with a weird sign I saw was. The first prompt worked with a stale location context and therefore was way off, the second prompt had to reach out to google servers, and came back with recognizing the physical space I was discussing, but told me that I was talking about an event that takes place in the museum next door that I had told the model was next door to the business in question, the third try it still seemed to understand where I was referencing, but insisted I couldn't possibly be talking about anything there.
It took 1 second on google maps to find exactly what I was referring to, which was the business in Google's system located at the exact map location the model had found.
I'm sick and tired of people turning to LLM and "AI" tools to pretend they are better, when the problem is that these companies don't even use existing good solutions because they just don't care.
thayne 10 hours ago [-]
One of the most common things I did with google assistant was tell it to remind me to do something, and it would reliably create a reminder on my calendar. With Gemini, it is very unreliable. Sometimes it does a google search. Sometimes it just opens a Gemini chat where it parrots back a (sometimes garbled) paraphrase of my request. Etc.
It's also really terrible at recognizing names of my contacts, probably because those names are not represented in the training data.
11 hours ago [-]
yu3zhou4 12 hours ago [-]
I hope you will find a solid solution for your father.
I was researching STT for people with speech disorders two years ago and essentially everything was boiling down to three problems at the end of the day - data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.
i mean for something this small, it can be fit into a l3 cache on a cpu and be essentially always on various purposes
Joel_Mckay 11 hours ago [-]
Have you tried playing music from his childhood for 20 minutes a day before his writing sessions?
In some cases, this may improve function for a few hours. Best regards =3
testycool 13 hours ago [-]
Unrelated: I love your username.
albert_e 13 hours ago [-]
What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.
raddan 12 hours ago [-]
What do streaming implementations do when a bigger context reveals a different interpretation/parse? When I use Whisper in the terminal, I can see it going back and correcting itself. Are corrections off the table for a true streaming transcription?
albert_e 4 hours ago [-]
I am guessing here.
Corrections based on larger context should also be part of the streaming output -- maybe include replacement text for previous chunk/s identified by chunk ID.
arbor-group 3 hours ago [-]
[flagged]
solarkraft 12 hours ago [-]
My mind is boggled by how many implementations miss this.
Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!
nicksaroha 11 hours ago [-]
Yeah! Streaming is crucial for any real time use.
NiloCK 6 hours ago [-]
My very first vibe-coded app (~Sonnet 3.6) was a dictation client for personal use.
I don't really understand how streaming would work compared against my normal flows. When I dictate, I set a toggle and then do stream of thought as I poke around between windows. When I'm ready to 'flush', I navigate to some target and give it focus for the text to flow.
Does dictation software now keep sort of unfocused floater previews and come with to-clipboard shortcuts or similar?
iforgotmypasswo 13 hours ago [-]
This is the main reason I lean on Deepgram over local services.
paynedigital 12 hours ago [-]
You can absolutely do high quality, low latency, even multilingual local streaming nowadays. As the commenter above says, Nemotron 3.5 Streaming is awesome. We make heavy use of it in our transcription app.
wkcheng 13 hours ago [-]
How does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.)
This definitely seems lighter and faster. How does accuracy compare?
jwr 12 hours ago [-]
People keep praising Parakeet, but I've found it to be worse than Whisper Large. Yes, it is much, much faster and smaller, but accuracy matters a lot if you are to use dictation regularly and seriously.
I ended up having AI optimize Whisper Large and create a plugin for TypeWhisper, and that's what I use (feeding the results through local Qwen 3.8 running under MTPLX).
wkcheng 9 hours ago [-]
When you are running the model locally and trying to do other tasks at the same time, Whisper Large feels too clunky on normal hardware.
I like Parakeet because it's good enough, relatively light, and fairly fast. I'm using this for things like meeting transcription and dictation. Since I'm sending most of the text to an LLM to clean up afterwards, it works well enough.
I'm hoping for something the size of Parakeet (or smaller) but better quality. It feels like with all of the advances in making smaller models better in the LLM space, someone should be able to come up with a lightweight and better quality speech to text model.
weitendorf 12 hours ago [-]
It’s really domain dependent I think. If you are doing anything conversational interfacing with less AI-familiar users, latency matters a lot.
If you’re feeding the results into a very smart LLM, it will figure out what you meant (but crucially ONLY if you warn it or tell it to do so, in some cases!). If you’re writing code directly or creating something for public consumption, you can’t tolerate mistakes. If you’re taking notes for yourself you just want it to work cheaply.
If you are ok with the complexity you can run both, and a Meta/Google open model with native audio, and let a smart LLM doctor it up. If performance really matters you can pay for a proprietary model or train one yourself. Until you get to that point, I think it probably doesn't matter much either way. Voice is just too easy to fiddle with
skolos 8 hours ago [-]
I confirm this as my experience as well. Parakeet did not perform well for me (English with heavy accent). Whisper large is much better. But once we get to model of this size look at qwen3-asr (1.7b parameters). I'm quite happy with how it performs (both transcription quality and speed on rtx5080). By the way the model is also available on openrouter and very/very cheap. Performance is reasonable (<1s response), but some requests are delayed (>10s)
grayrest 8 hours ago [-]
For English I've found Parakeet v2 and Parakeet unified to perform significantly better than v3. For deposition audio in a quiet room and neutral american accents I've hit 96-97% on a number of sessions with v2 and I've seen as low as 83% for heavy accents with overlapping speakers.
jhatemyjob 4 hours ago [-]
Whisper large-v3-turbo is still the SOTA non-cloud model? Damn. I'm wanting to move away from it to something smaller/better but based on this comment it seems like I made the right move by sticking with whiser.cpp this whole time. It's been 2 years since large-v3-turbo was released...
theturtletalks 13 hours ago [-]
Parakeet is the gold standard. With models like moonshine and koroko (TTS model), it’s more about embedding the model in the application itself. If you’re using Parakeet, embedding it in the application is not feasible.
I use parakeet with superwhisper, and I’m making another app that has SST and TTS built in, and I want to use my downloaded parakeet model, but it seems there’s so many different implementations from ONNX to whisper, it’s not easy to use your downloaded models. So models like moonshine and this one allow you to just embed it into your application simply. It might not be as good as parakeet, but it gets you 80% of the way there.
zimpenfish 13 hours ago [-]
Tried it on a random TV episode and it seems to get stuck sometimes where it just outputs "Thank you." as a default - at one point emitting that for 60s of dialogue (and no, the episode does not have someone repeating "Thank you." for 60s.) Happens several times during the transcription.
jwr 12 hours ago [-]
It's funny how so many models tend to generate "Thank you. Don't forget to subscribe" or "Thanks for watching" if there is silence. Shows you what they've been trained on :-)
aqfamnzc 9 hours ago [-]
Reminds me of "foreign" and "[Applause]" frequently appearing on YT auto-subtitles.
nl 5 hours ago [-]
I recently built a 3D printed ESP32 based transcription device[1] that offloads transcription to Parakeet running in the browser connected to the ESP32. I haven't done my own measurements of error rate but it'd be interesting to see how well this runs on the ESP32 itself.
In my case I'm already using one core to run DSP for a beamforming mic array (which works amazingly well for noise cancellation!) so I don't have huge amounts of free processing though.
Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!
I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!
flowerlad 9 hours ago [-]
Apple needs to incorporate this into iOS ASAP. This works much better than the speech recognition in iOS when you use technical terms. One of the most frustrating parts of iOS is speech-to-text in iMessage. For me no feature is more important in a phone.
Try this example: My website uses ASP.NET technology and I am using .NET 10.0. Works perfectly in Whistle, but not in iOS.
joewhale 14 hours ago [-]
I initially read this as whistle to text, which would be way cooler.
Did anyone ever read The First Circle by Solzhenitsyn? It is amazing: in a Stalin gulag in the 1950's USSR, a bunch of Soviet engineers / techies are tasked with creating a speech recognition system. They show fake progress in order to keep from being jailed in worse ways. It is autobiographical, so good, but I read it so long ago and it convinced me how hard speech recognition would be, maybe even impossible!
Are there such STT models that allow ‘context’ to be provided? I’d love to be able to dictate for programming sessions. The problem is that STT models don’t handle jargon and acronyms well, especially if they’re specific to the project or company. Is there like a type of model you can provide a custom vocabulary to without needing training?
RexHuang 3 hours ago [-]
Nice work. One gap though: no Chinese in the seven languages — and Chinese is where a lot of the on-device demand actually is. Any plans?
nine_k 2 hours ago [-]
Chinese has a huge amount of homonyms, not differing in their pronunciation in any way at all, except in their contextual meaning. In languages like Spanish it's almost enough to map he sound to a written word, and there'd be really few places where the written rendering would be ambiguous. In English it's worse, you need a larger model of the language to e.g. correctly choose between "no" and "know", or "there" and "their", depending on the context.
But Chinese is in another league; the same spoken word may have five meanings, or ten meanings (open zhongwen.com and check out), and you have to build the complete sentence as you parse the sounds, asses its meaning (or several possible meanings, maybe in the context of a few previous sentences), and choose the written word that would match the meaning. You need to carry a lot larger "sense-making" model along with your phonetic, grammatical and syntactic models.
blagui 5 hours ago [-]
Error rate is crazy high.
And comparing to low models and omit a lot of other bigger open models.
amelius 10 hours ago [-]
Let's say I want to build a hardware product now, voice-controlled, with voice feedback, so STT, LLM, and TTS. All local. What are the best libraries to do this now, say with 8GB of GPU memory available?
system2 8 hours ago [-]
Whisper local (openai). Works incredibly well; even translates.
HenryNdubuaku 8 hours ago [-]
thats what whistle is for :)
yjftsjthsd-h 7 hours ago [-]
It might work, but 8G of GPU is wild overkill for whistle.
aavisangle 3 hours ago [-]
I think it works great on Mac but lags on Windows. The accuracy on Windows is lower than Mac.
Can't even be bothered find a human to write your model announcement article huh.
e12e 13 hours ago [-]
Hm. I saw language=detect and tried some Japanese - which (given the actual list of supported languages) unsurprisingly turned into some mangled Spanish.
Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning).
So, reasonable, but limited?
11 hours ago [-]
HenryNdubuaku 8 hours ago [-]
it does not support those languages, its explained in the writeup
e12e 7 hours ago [-]
Maybe I was unclear - I tried Japanese before I saw the list, then English and with that sample of one, it got it wrong.
So limited selection of languages, not stellar performance?
TomGarden 11 hours ago [-]
The Qwen STT model that's leading the open weight leaderboard right now is excellent. Parakeet v2 is blazing fast and accurate enough on English. I feel like STT gains from this point on will be marginal, especially given that you can do a quick LLM pass afterwards with a small model
MisterMunchkin 12 hours ago [-]
For something you could plausibly ship inside a webapp, it’s very good. I can definitely think of some cool uses for this.
armcat 14 hours ago [-]
Those are insane benchmarks at this size. Well done!
flytomoon 7 hours ago [-]
Thanks for putting in the work. Only a matter of time before all of this runs locally at very high accuracy. And can handle complex vocabularies.
kamranjon 14 hours ago [-]
Sooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good.
est 5 hours ago [-]
I really hope there's a TTS model in 16.9MB
Kokoro is good, but can't change the pitch easily.
rafaelm 13 hours ago [-]
Huh, this was really confusing. I already had an STT app called Whistle on my Android phone.
jakobov 11 hours ago [-]
ZWhispr seems to be the best for those who care about accuracy as it uses three SOTA models.
I find the whole category of "added value" STT apps funny at this point.
Built my own in a day that blows them out of the water. You can quite literally pick any sufficiently good local model or API provider and combine it with Cerebras for cheap and very fast AI post processing and formatting.
jakobov 7 hours ago [-]
If you want the best accuracy, then this is not true. Combining models will beat any single model. There's also a lot of work that goes into optimizing dictionary words and screen context, and knowing which ones to, to add to model context.
Lebenita 10 hours ago [-]
macOS only (like so many in this space)
jakobov 7 hours ago [-]
What operating system are you hoping for?
cellular 5 hours ago [-]
Can it run on any of the cheap uC / esp / stm ?
sheephess44 4 hours ago [-]
[flagged]
jayshah5696 13 hours ago [-]
This is actually a really great release. Congratulations team. I just tried few words. My Indian accent also was able to pick up.I'm gonna run it on my Linux Box.
properbrew 12 hours ago [-]
Might look into embedding Whistle into Whistle if it can make it 30x smaller (https://play.google.com/store/apps/details?id=com.blazingban...) - It's a shame there isn't as many languages supported though, I'm surprised at the amount of non-english downloads (I really shouldn't be, of course non-english speakers want dictation) of the app there is.
rshemet 12 hours ago [-]
hey, Roman here from Cactus, thank you for the feature!
opening this thread for questions/feedback if you have any
villgax 3 hours ago [-]
Actually bad WER
mrkn1 14 hours ago [-]
love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages. Unlimited transcription for free.
Initial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly.
MayeulC 13 hours ago [-]
Speech to text I assume? Maybe it has to do with your a accent or pronunciation? You could contribute a bit to Mozilla's Common voice, if that's the case. I assume it is part of every STT training corpus.
saturn8601 13 hours ago [-]
Yes sorry Speech to Text. I have a standard US East coast accent but sometimes I speak a little mumbly. When I made an effort to speak more slowly and with a cleared throat there was some improvement but still not writing all words.
pzo 12 hours ago [-]
tested in polish and unless you speak very loud, clear and slow is not that good, parakeet definitely better.
aidotguru 14 hours ago [-]
eager to see if working in android phones
rpdillon 13 hours ago [-]
FUTO keyboard (open-source, free) runs entirely on-device and has extremely good STT accurary, especially with the 70M parameter model. I've used it for years now and love it.
Edit: As others have pointed out, this is not actually open source. It's source-available, which is quite a bit different because folks can't fork and distribute it as easily. The license also appears to be revocable and non-transferable, which makes it different from open source licenses.
robertlane0 13 hours ago [-]
I'd classify it as "source-available" given the noncommercial clause in the license.
Good call out. I tend to be very sensitive about these things, and this is a case where I messed up. Thanks for the correction.
lrvick 13 hours ago [-]
Futo only produces source-available proprietary software. They most certainly are not Open Source, though they unfortunately lied about this a lot before they got called out enough times.
Good call... I got that one wrong. Thanks for the pointer!
helterskelter2 13 hours ago [-]
I've used FUDO keyboard for a long time, but I never tried using the voice input, so I'm testing it now. Let's see how well it transcribes everything.
...Okay that was pretty good.
tecleandor 14 hours ago [-]
Spanish is not good (seems to write non existing words and/or with terrible typos...) but English seem to work good even with my (Spanish) accent...
kaoD 14 hours ago [-]
Spanish from where? Here (Castilian Spanish) it seemed to work fine.
tecleandor 13 hours ago [-]
Madrid. But it will only work properly if I'm clearly dictating with a very regular rhythm (ViaVoice dictation, if anyone remembers...). If I use a more natural/conversational rhythm (no slang, no abbreviations...) it easily confuses words.
chilicuil 13 hours ago [-]
Mexican and venezuelan aren't detected correctly
snvzz 6 hours ago [-]
Ships with a RISC-V build. This is much appreciated!
12 hours ago [-]
lab14 10 hours ago [-]
Tried it a few times with English, Spanish and French and the quality/accuracy is pretty "meh". If the model doesn't really work, it doesn't matter if it fits in 1MB.
contingencies 11 hours ago [-]
For speech to text UX I currently use https://handy.computer/ as it's cross platform and open source. With that I am currently using Parakeet Unified EN 0.6B and finding it excellent. Often I use it to talk to AIs without giving them audio, which works very well. Honestly, I would never go back to typing now. Promised since ~Y2K, the tech is finally here. You really notice it when you wake up at 2AM and don't want to wake people ... it can get really annoying reverting to key-tapping. My long-gnawing fear of losing my hands to RSI is no longer a thing, and I can focus on losing them to another hobby: like sailing or machining! Just bought a band saw...
try-working 12 hours ago [-]
I built an STT plugin for DeepSeek Harness that uses this Whistle model as well as a larger one from Desert Ant Labs: https://github.com/try-works/dsh-stt
- my whistle setup (tested can run on echo show with <1s response): 206 correct out of 208
These were not tested on echo show yet - on my pc for now:
- Vosk small, phrase grammar: 184/208
- Speech-to-Phrase 1.4.3 (Kaldi): 118/208
- suggested sphinx: 63/208 and on some examples took 13s on i7 14700k pc
looks like my customized whistle works better for me than these alternatives. But more testing wouldn't hurt.
XML Web Services were awesome, but the SOAP and XML (where the hard parts that should have been given some batteries included defaults) I think was too much of a boat anchor to overcome.
And then the ruby/rails wave came and made 'rest' JSON the Silicon Valley hearthrob, and WSDLs were out like yesterdays trash.
Are they still used in businesses IRL? Yes! I worked for one that has probably (probably) moved off of them by now, but as recently as 2023 there were still some bank-facing power with SOAP RPC .asmx endpoint servers.
I follow hypermedia.systems and htmx/datastar/alpine.js because I'm still trying to get back that powerful _web_ tech advancement rather than squeezing everything into the javascript client side workaround. Typescript is great, a good poor man's F# and leagues ahead of ES3 (which I dabbled in / torture LLMs with so my old retrocomputers can still do web things), but its still bringing along so much baggage that could be lighter weight for older devices AND fits the tech utopianist promise of the original web (which is half of why I play with computer stuff).
By law, every single government form you're able to file is supposed to have an XML Schema available in a centralized registry. This regulation is widely ignored in practice, particularly by local / municipal governments.
This used to be more important in the past, as such a form could automatically / semi-automatically get an entry in the EPUAP form catalog and be made electronically fillable (EPUAP being the now-deprecated centralized government bureaucracy portal basically). As far as I understand, the way that worked was through XSLT and XML Forms. There was some weirdness about each document having a fillable / form view and a preview, I think XML Forms was somehow used to generate the XHTML form, while the XSLT sheet could only generate an XHTML preview of a complete, filled-in version. There were also some custom annotations for auto-fill and such, that was partly done by most documents relying on standardized schemas for entities like "person", "address" or "company".
Since we moved to E-Deliveries and lost a central place for these forms to live in, this is (AFAIK) a bit less important and less common, but internally, things are still XML. If you're filing something like an ID renewal application, even through a newer, more user-friendly frontend, it's still just XML underneath. If you sign something through podpis.gov.pl (the standardized e-signature solution for government paperwork), you can even download that underlying document, both signed and unsigned, and see what that XML is. The pre-signing document preview generated by that site still comes from the XSLT I believe. The signatures themselves are, unsurprisingly, also done via the XML signatures spec.
Incidentally, the European E-Delivery system itself also relies on XML, WSDL and Soap pretty heavily. For those unfamiliar and/or not in Europe, it's basically "email but for the government", with all the guarantees and legal obligations of physical mail, cryptographically-attested confirmation of receipt, proper identity verification and assurance, cross-provider address portability, deployed to a lesser or greater extend in many EU countries and set to replace physical mail.
Much larger though, I think I went full precision and its around 2GB.
I chuckled at this because my inner voice had an accent as I was reading your comment, due to your writing style.
With a restricted grammar, built in Windows voice recognition, all on device, has managed this exact use case quite well for over a decade. I used it to try and build a clone of the various paid apps that allow you to issue orders to Arma soldiers with voice commands
I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.
But every single sound he makes with his mouth ends up on the page too.
Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...
You can set it "Push to talk" mode (like a walkie-talkie radio), and when you're done talking and release the button, it can paste the text into any text field.
You can even replicate ChatGPT voice conversation mode, by having Handy as your speech input, and then (I forgot the extension) enabling a speech-to-text model for OpenCode. Surprisingly relaxing flow for certain tasks, like tweaking a website's styles.
I am especially interested in this part. Could you please share the prompt you are using to instruct the LLM to clean up the dictation? Thanks in advance!
I prompt it to:
"Attached (or underneath) is the transcript of a self recording i've done with tons of rambling and some incorrect words transcriptions, please do a pass clearing out and arranging any typos or possible misunderstandings. Keep original in parenthesis when not sure if it's a misunderstanding. Do not summarize or alter the nature of the content, simply tidy the transcript."
In case it's helpful to anyone else using it, at first it felt a bit slow to me, because there was a noticeable pause after I finished a message before it would quickly type it all out. I changed the input method from direct to clipboard and it's way faster now, almost instantaneous.
Because I have seemingly mixed opinions on it, on one hand, I did put the effort but on the other, the output is AI generated so I am unsure about sharing it with others (because they might think its AI generated)
Do you use it for very small edits (removing just the uhhm's?) or for slightly more edits.
The way that I use it sometimes is that while thinking, I will write something which can sometimes make me feel as if a better re-write can better explain my thoughts or rephrasing it as such. For example. I will think about X topic, connect it to Y, then try to add some more points about X again.
I found LLM's to do a really decent job at generating the final outputs as such, but as I said, I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated and the end user not knowing if I put an actual effort into creation of it or not.
Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it.
"what are your observations on feeling as if sharing that output though?"
and "Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it."
could you rephrase the question
for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.
Yes, for doing extremely mild edits (like just removing uhh's etc.) this might be true
but I sometimes feel as if writing can allow me to shift paragraphs, so if I am writing para 1, para 2, I can shift back to para 1 and write another sentence in it and edit some parts of para 1 to include that point
but when I am doing STT, although I can move towards the other para, I find myself just speaking in a complete flow and just write in para 2 only.
Thus when I ask AI to write, I would prefer it to move the statements to appropriate paragraphs and in just general, create a more comprehensible viewpoint from all the STT text that I had written.
This does generate AI generated text which can be detected as such. Uploading it on blogs makes me feel as if people might read what they might consider "AI slop" and so the ethics part (as I myself don't wish to read AI slop)
the problem with AI written or edited texts is that I am unsure of how much effort the other person has put in (just a single prompt or a detailed thought was put in), and I feel as if, others feel the same way.
Should one try to show the rough draft as well to try to show that it was an effort which was human generated or that human effort was used, but that means having a proper disclosure that it was AI-generated/AI-assisted, which I feel as if offputs a lot people (including me) as because of the above logic, that there's still friction for the user within testing if real effort was put into place and I am unsure how effective sharing drafts of it could be.
I don't want my blogs to be tainted and treated as AI-slop because I care about them so I am unsure of what to do. I have multiple things that I have written which if I pass through AI can create some meaningful blog piece but as it stands, they are rough drafts and I find myself putting low efforts or being lazy in actually editing them myself as well (and potentially putting in multiple hours) when AI can be used to help create a more polished version just as well and get across my point.
The usecase for small models like this is making on-device STT/TTS more accessable. This is important if your usecase is sensitive to either privacy or latency, but this comes at the cost of quality.
My experience has been that these small TTS models are unexpectedly good if your audio is in distribution (western accents, higher quality audio, common vocabulary), but pretty quickly degrade as you move outside of that. They often dont support more complex features such as diarization, multilingual, or realtime streaming either.
It's a very compelling aspect of the problem.
If you can get a model under certain size thresholds, that means you can eliminate latency domains. For example, if the model is able to fit entirely inside L3, the latency of servicing requests drops by an order of magnitude (or better) compared with a model that resides in L3+DRAM.
And even when it's getting it right, and if there's only one Walmart in Beaverton, it still to this day needs to ask "One option is Walmart on Expressway Road in Beaverton...." Maybe it's correct in its 0% confidence level there, since it's so bad, but... I don't get how you could design something that bad, even before LLMs existed. I feel like I could do better, even using their Speech-to-text engine, with the processing backend built of pure regexes and if/elses.
"Hey Siri, play [song]"
Leads to, take your pick:
- "You'll need to unlock your iPhone first."
- "I couldn't find [song] on Podcasts" (??????)
- "Playing [a totally different song]"
- "I couldn't find any music by [song, but it thinks it's a band]"
- "Playing music by [song, again it thinks it's a band]"
And don't forget whatever the current phrasing is for "I'm sorry, my shit's all fucked up" and "My network connectivity had a blip and I'm unwilling to retry" and "Even though I have on-device STT models, and now LLMs too, and an on-device database of your music, which is downloaded, I won't bother without the cloud.
but once you calm down and stop hyperventilating from my suggestion, you'll see that the only reasonable and pragmatic course of action is to move to android and linux.
If you can do something with an extremely limited vocab, voice recognition was fine using off the shelf microchips in the 70s, where you wired in a microphone connection and had discrete pins for output actions.
LLMs are basically only useful for utterly free form transcription, but that doesn't actually help you turn that into tasks to perform and parameters for those tasks
The core "problem" in voice recognition is that freeform speech is an abysmal UX paradigm and provides zero discoverability, and LLMs IMO have not improved the situation of actually doing anything with the resulting text.
The other day I tried to prompt Gemini 3 times to tell me what the heck the business with a weird sign I saw was. The first prompt worked with a stale location context and therefore was way off, the second prompt had to reach out to google servers, and came back with recognizing the physical space I was discussing, but told me that I was talking about an event that takes place in the museum next door that I had told the model was next door to the business in question, the third try it still seemed to understand where I was referencing, but insisted I couldn't possibly be talking about anything there.
It took 1 second on google maps to find exactly what I was referring to, which was the business in Google's system located at the exact map location the model had found.
I'm sick and tired of people turning to LLM and "AI" tools to pretend they are better, when the problem is that these companies don't even use existing good solutions because they just don't care.
It's also really terrible at recognizing names of my contacts, probably because those names are not represented in the training data.
I was researching STT for people with speech disorders two years ago and essentially everything was boiling down to three problems at the end of the day - data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.
In some cases, this may improve function for a few hours. Best regards =3
Corrections based on larger context should also be part of the streaming output -- maybe include replacement text for previous chunk/s identified by chunk ID.
Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!
I don't really understand how streaming would work compared against my normal flows. When I dictate, I set a toggle and then do stream of thought as I poke around between windows. When I'm ready to 'flush', I navigate to some target and give it focus for the text to flow.
Does dictation software now keep sort of unfocused floater previews and come with to-clipboard shortcuts or similar?
This definitely seems lighter and faster. How does accuracy compare?
I ended up having AI optimize Whisper Large and create a plugin for TypeWhisper, and that's what I use (feeding the results through local Qwen 3.8 running under MTPLX).
I like Parakeet because it's good enough, relatively light, and fairly fast. I'm using this for things like meeting transcription and dictation. Since I'm sending most of the text to an LLM to clean up afterwards, it works well enough.
I'm hoping for something the size of Parakeet (or smaller) but better quality. It feels like with all of the advances in making smaller models better in the LLM space, someone should be able to come up with a lightweight and better quality speech to text model.
If you’re feeding the results into a very smart LLM, it will figure out what you meant (but crucially ONLY if you warn it or tell it to do so, in some cases!). If you’re writing code directly or creating something for public consumption, you can’t tolerate mistakes. If you’re taking notes for yourself you just want it to work cheaply.
If you are ok with the complexity you can run both, and a Meta/Google open model with native audio, and let a smart LLM doctor it up. If performance really matters you can pay for a proprietary model or train one yourself. Until you get to that point, I think it probably doesn't matter much either way. Voice is just too easy to fiddle with
I use parakeet with superwhisper, and I’m making another app that has SST and TTS built in, and I want to use my downloaded parakeet model, but it seems there’s so many different implementations from ONNX to whisper, it’s not easy to use your downloaded models. So models like moonshine and this one allow you to just embed it into your application simply. It might not be as good as parakeet, but it gets you 80% of the way there.
In my case I'm already using one core to run DSP for a beamforming mic array (which works amazingly well for noise cancellation!) so I don't have huge amounts of free processing though.
[1] https://github.com/nlothian/earwing
I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!
Try this example: My website uses ASP.NET technology and I am using .NET 10.0. Works perfectly in Whistle, but not in iOS.
https://en.wikipedia.org/wiki/Silbo_Gomero
[1]: https://fr.wikipedia.org/wiki/Langage_siffl%C3%A9_d%27Aas
[2]: https://www.dailymotion.com/video/x4lxhsi
Highly recommended. I see there is a "new" (2009, I getting old) edition & translation - https://www.amazon.com/First-Circle-Aleksandr-I-Solzhenitsyn...
But Chinese is in another league; the same spoken word may have five meanings, or ten meanings (open zhongwen.com and check out), and you have to build the complete sentence as you parse the sounds, asses its meaning (or several possible meanings, maybe in the context of a few previous sentences), and choose the written word that would match the meaning. You need to carry a lot larger "sense-making" model along with your phonetic, grammatical and syntactic models.
Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning).
So, reasonable, but limited?
So limited selection of languages, not stellar performance?
Kokoro is good, but can't change the pitch easily.
https://zwhispr.com/
Built my own in a day that blows them out of the water. You can quite literally pick any sufficiently good local model or API provider and combine it with Cerebras for cheap and very fast AI post processing and formatting.
opening this thread for questions/feedback if you have any
https://futo.tech/
Edit: As others have pointed out, this is not actually open source. It's source-available, which is quite a bit different because folks can't fork and distribute it as easily. The license also appears to be revocable and non-transferable, which makes it different from open source licenses.
https://github.com/futo-org/android-keyboard/blob/master/LIC...
https://github.com/futo-org/voice-input/blob/master/LICENSE....
...Okay that was pretty good.