How it explains each slide
A slide is not running text: it is loose blocks on a page, and often a chart, a table or a diagram. Filmina works on that, and from there come the two things that separate it from a text-to-speech reader: it explains instead of reading, and it can point at what it is talking about.
What gets read off the slide
Two things are stored for each page: the image, exactly as it looks, and the text blocks with their coordinates. Those coordinates are what later lets the frame surround the exact block the voice is talking about. The image goes to the model along with the text, so it can also read what is not written out in blocks.
Not everything written on the page is content. Before explaining, running headers, footers, page numbers, dates and stray web addresses are discarded: page numbers and dates are recognized by their shape, and headers and footers because they repeat page after page along the top or bottom edge. If they got through, the total coverage rule would force it to explain the footer to you.
A scanned PDF is a photo of each page with no text inside, so there are no blocks to read or point at. Instead of making something up, Mina tells you when you upload it. If you have the original deck, upload that instead.
The segments
The explanation does not come back as one block of text: it comes back split into segments of one idea each, between three and seven per slide, ten to thirty seconds spoken each. Every segment knows which blocks of the page it covers, and only groups blocks that sit together: a heading with its paragraph, a box with its list.
That is what lets the frame move: the segment being spoken says where to look. It is also the unit you navigate with, because the up and down arrows jump from one idea to the next, not from one second to the next.
What the model is not allowed to do
The explanation is written to be heard, not read, and that comes down to concrete rules:
- No reading verbatim. It reformulates, connects and says what each thing means and why it matters. The slide is already in front of you.
- Absolute precision. Figures, names, formulas, dates and technical terms are kept exactly as they appear. Nothing is rounded and no term is swapped for an informal synonym.
- No preamble and no wrap-up. None of "let's take a look", "in this slide we see" or "we are on 22 of 62". Those are the seconds you spend waiting for the part that matters, so the first sentence is already asserting something about the topic.
- No markdown. Running prose: no asterisks, no bullets, no numbering, no tables, no emoji. None of that has a sound.
- Notation in words. "P of A given B", not "pee paren a bar bee". Symbols and formulas are written the way they are said before they reach the voice, which is exactly where text-to-speech readers break, and the symbol said is the one on the slide, not the one the subject usually uses.
- What was already said is not repeated. If a slide carries the same heading as the previous one, what that heading means has already been explained: it is not defined again, not even in passing, and only what the new slide adds gets told.
Coverage is verified in code
Asking a model to skip nothing is not enough: it has to be checked. When the answer comes back every block is reviewed to make sure none was left unexplained, and if one is missing it is asked for again on its own. The pictures on the slide go into the same count: each one gets explained, or the model has to say it is only decoration. And if a text block is still missing after the second request, it is read out as it stands: less elegant, but not left out.
Coverage does not depend on the model complying. It is checked, and what is missing is asked for again.
Charts, tables and diagrams
The model gets the slide as an image, the same one you are looking at, on top of its text. That is how it reads what the blocks do not hold: the data in a chart or a table, the steps of a diagram, a screenshot, the small note in a corner. A table is told with its numbers, not just announced. And whatever it cannot read with certainty, it does not guess.
When a segment is only about a picture, the frame goes to it. A slide that is nothing but an image gets explained too, if the image says something. If it only decorates (a logo, a background, a photo that comes along), the voice asks you to look at it and moves on, instead of describing an ornament just to have something to say.
A diagram's loose labels are a separate case. They get explained together, in a single segment, describing what the diagram shows and how those elements relate according to the slide. They are labels on a drawing, not claims made by the material: giving each loose word its own segment would mean inventing a definition for each one. Formulas, on the other hand, are not labels: for each one it explains what it computes and what each term stands for.
Your notes and your material count too
If the document has class notes, each slide picks up the lines from your notes that share the lecture's less common words with it, and the model gets them along with the blocks. What the note adds gets told; what only underlines something already on the slide gives that part more weight. How to add notes and how to leave them out is in class notes.
The same goes for material you added to the class, like a book, a handout, a recording or a video: the passages that cover the same ground as the slide come along with it. They are only used where they make the slide clearer (a fuller definition, an example, a reason why), and if the material contradicts the slide, the slide wins.
Which language it explains in
The explanation is in the language the deck is written in, Spanish or English, not the app's: a lecture in English gets explained in English even if Filmina is set to Spanish. It is checked once, on the first pages, and the voices on offer are the ones for that language. If there is not enough text to tell, the app's language is used.
If something slips through
The rules make the model invent far less, not never. That is why every paragraph in the script has a "This is wrong" button: tap it and Filmina looks at the slide again, with its image, your notes and your material, and if it agrees with you it rewrites that paragraph and voices it again. What gets confirmed is remembered for that lecture. How it works is in the viewer.
Prepared ahead
While you listen to one slide, the next two are being prepared in the background, so moving on rarely means waiting. The whole document is not prepared in one go: only what you are about to hear.
The slide you are looking at always goes first: if you jump to another one, it moves ahead of the ones being prepared. Closing the document stops whatever was still queued, and when you come back it picks up what is missing.
What gets generated is kept in your account. The script is saved per language and the audio per voice: changing voice does not rethink the explanation, it just says it again. And listening again to something already explained does not generate it again or use up minutes from your plan.