How it explains each slide

A slide is not running text: it is loose blocks on a page, and often a chart, a table or a diagram. Filmina works on that, and from there come the two things that separate it from a text-to-speech reader: it explains instead of reading, and it can point at what it is talking about.

What gets read off the slide

Two things are stored for each page: the image, exactly as it looks, and the text blocks with their coordinates. Those coordinates are what later lets the frame surround the exact block the voice is talking about. The image goes to the model along with the text, so it can also read what is not written out in blocks.

Not everything written on the page is content. Before explaining, running headers, footers, page numbers, dates and stray web addresses are discarded: page numbers and dates are recognized by their shape, and headers and footers because they repeat page after page along the top or bottom edge. If they got through, the total coverage rule would force it to explain the footer to you.

A scanned PDF is a photo of each page with no text inside, so there are no blocks to read or point at. Instead of making something up, Mina tells you when you upload it. If you have the original deck, upload that instead.

The segments

The explanation does not come back as one block of text: it comes back split into segments of one idea each, between three and seven per slide, ten to thirty seconds spoken each. Every segment knows which blocks of the page it covers, and only groups blocks that sit together: a heading with its paragraph, a box with its list.

That is what lets the frame move: the segment being spoken says where to look. It is also the unit you navigate with, because the up and down arrows jump from one idea to the next, not from one second to the next.

What the model is not allowed to do

The explanation is written to be heard, not read, and that comes down to concrete rules:

Coverage is verified in code

Asking a model to skip nothing is not enough: it has to be checked. When the answer comes back every block is reviewed to make sure none was left unexplained, and if one is missing it is asked for again on its own. The pictures on the slide go into the same count: each one gets explained, or the model has to say it is only decoration. And if a text block is still missing after the second request, it is read out as it stands: less elegant, but not left out.

Coverage does not depend on the model complying. It is checked, and what is missing is asked for again.

Charts, tables and diagrams

The model gets the slide as an image, the same one you are looking at, on top of its text. That is how it reads what the blocks do not hold: the data in a chart or a table, the steps of a diagram, a screenshot, the small note in a corner. A table is told with its numbers, not just announced. And whatever it cannot read with certainty, it does not guess.

When a segment is only about a picture, the frame goes to it. A slide that is nothing but an image gets explained too, if the image says something. If it only decorates (a logo, a background, a photo that comes along), the voice asks you to look at it and moves on, instead of describing an ornament just to have something to say.

A diagram's loose labels are a separate case. They get explained together, in a single segment, describing what the diagram shows and how those elements relate according to the slide. They are labels on a drawing, not claims made by the material: giving each loose word its own segment would mean inventing a definition for each one. Formulas, on the other hand, are not labels: for each one it explains what it computes and what each term stands for.

Your notes and your material count too

If the document has class notes, each slide picks up the lines from your notes that share the lecture's less common words with it, and the model gets them along with the blocks. What the note adds gets told; what only underlines something already on the slide gives that part more weight. How to add notes and how to leave them out is in class notes.

The same goes for material you added to the class, like a book, a handout, a recording or a video: the passages that cover the same ground as the slide come along with it. They are only used where they make the slide clearer (a fuller definition, an example, a reason why), and if the material contradicts the slide, the slide wins.

Which language it explains in

The explanation is in the language the deck is written in, Spanish or English, not the app's: a lecture in English gets explained in English even if Filmina is set to Spanish. It is checked once, on the first pages, and the voices on offer are the ones for that language. If there is not enough text to tell, the app's language is used.

If something slips through

The rules make the model invent far less, not never. That is why every paragraph in the script has a "This is wrong" button: tap it and Filmina looks at the slide again, with its image, your notes and your material, and if it agrees with you it rewrites that paragraph and voices it again. What gets confirmed is remembered for that lecture. How it works is in the viewer.

Prepared ahead

While you listen to one slide, the next two are being prepared in the background, so moving on rarely means waiting. The whole document is not prepared in one go: only what you are about to hear.

The slide you are looking at always goes first: if you jump to another one, it moves ahead of the ones being prepared. Closing the document stops whatever was still queued, and when you come back it picks up what is missing.

What gets generated is kept in your account. The script is saved per language and the audio per voice: changing voice does not rethink the explanation, it just says it again. And listening again to something already explained does not generate it again or use up minutes from your plan.