The voice that explains

The voice is local too. It is synthesized by Kokoro, an 82-million parameter model that ships inside the app and works against your Mac's GPU: there is no voice service on the other end, so what gets said never leaves your machine and no connection is needed to hear it.

The voices

Voices belong to a language, not to the app: a Spanish voice reading English sounds like someone reading in a language they do not speak. Depending on the language you are listening in, there are three:

You change the voice from the dock settings, at any time. Changing it does not rethink the explanation: the script is already written and what gets redone is the saying of it.

The speed

From 0.8× to 1.75×, in seven steps. Speed regenerates nothing: it is the audio that already exists, said faster, so you can move it mid-sentence with no wait.

Straight through

When the slide's last idea ends, and the question if questions are on, it moves to the next one and keeps going. That is what turns Filmina into something you listen to straight through instead of something you request one slide at a time. To stay on a slide, pause.

Why they do not collide

The voice that speaks and the ear that listens to you both use the GPU, and on macOS, both at once, they kill the process. So they take turns: while one works, the other waits. What you lose is a couple of seconds; what you gain is the app not closing in the middle of a lecture.