Skip to main content

Extract Book References From YouTube Videos: The References & Key Data Recall Style

Long videos and podcasts bury the books, studies and statistics they mention. How the References & Key Data Recall style pulls them back out.

You remember that they mentioned a great book. You do not remember which one.

Ninety minutes into an interview, a guest says the thing that reframed the whole problem for them was a particular book. You are driving, or running, or doing the washing up. You do not write it down. Two weeks later all that survives is the shape of the memory: someone recommended something good, on that episode, somewhere in the middle.

This is a retrieval problem, not a comprehension problem. You understood it perfectly at the time. What you lost was the one concrete, searchable detail — the title, the surname, the exact figure — that would let you find it again.

Podhoc now ships a podcast style built for exactly that: References & Key Data Recall. This article explains what it extracts, what it genuinely cannot do, and how to use it.


Why long-form audio buries its own references

Three things conspire here, and none of them is your memory failing.

There is an enormous amount of this material, and people consume it in situations where writing is impossible. YouTube reported more than 1 billion monthly active viewers of podcast content in February 2025, and that “viewers watched over 400M hours of podcasts monthly on living room devices.” Living rooms, cars, gyms and kitchens are precisely the places where you cannot pause to take a note.

The episodes are long and the references are unsignposted. The Spotify Podcast Dataset (Clifton et al., 2020) assembled “approximately 100K podcast episodes” amounting to “over 47,000 hours of transcribed audio” — an average in the region of half an hour per episode across a broad sample, and interview shows routinely run three or four times that. A two-hour conversation can name a dozen books, several researchers, a handful of tools and a scattering of statistics, none of them announced, none of them collected anywhere at the end.

Speech leaves nothing to search. An article gives you Ctrl-F and a bibliography. Audio gives you a scrub bar. Even when an auto-generated transcript exists, finding a name in it requires you to already know how that name is spelled — which is the thing you are trying to recover.


What the References & Key Data Recall style produces

The style does one job: it goes through your sources and recaps the concrete artefacts named in them — the books and their authors, the studies and papers, the tools and products, the statistics and figures, and the other key data points — together with the context of why each one came up.

It is deliberately not a summary. Deep Dive gives you the conversation. Simplified Explanation gives you the argument stripped to its bones. Critique interrogates the reasoning. References & Key Data Recall gives you the bibliography — the part every other format quite reasonably leaves out.

It defaults to a single narrator, because a list read aloud does not benefit from a second voice. You can browse the full lineup on the audio styles page.


What it cannot do, stated plainly

Two limits. Both matter more than any feature on the list above.

It can only surface what is actually present in the source. If a guest says “there is a brilliant book on this by someone at Stanford” and never names it, nothing recovers the title — the style extracts, it does not research. The same applies to half-remembered statistics: “something like a third of people” stays “something like a third of people”. The output is only ever as specific as the speaker was.

Spoken proper nouns are the hardest thing in the whole pipeline. A YouTube video reaches Podhoc as a transcript, and machine transcription of names is measurably worse than transcription of ordinary words. The Earnings-21 benchmark (Del Rio et al., Interspeech 2021) was built specifically to test this on “entity-dense speech” and concluded that “ASR accuracy for certain NER categories is poor, presenting a significant impediment to transcript comprehension and usage.”

It does not stop at transcription, either. An ACL 2023 study, “Why Aren’t We NER Yet? Artifacts of ASR Errors in Named Entity Recognition in Spontaneous Speech Transcripts” (Szymański et al.), reported that entity-recognition “models fail spectacularly even if no word errors are introduced by the ASR.” Unscripted speech alone — false starts, missing punctuation, no capitalisation — is enough to degrade entity extraction before a single word is mis-heard.

So the honest framing is this: the output is a lead list, not a citation list. An unfamiliar surname may come back mis-spelled. A foreign-language title may be anglicised. A study may be pinned to the year it was discussed rather than the year it was published. Every one of those is a five-second fix once you have the lead — and unrecoverable if you have nothing. But verify before you cite, and verify before you buy.


How to select the style

The flow is the same in the web app and the mobile app.

  1. Start a new podcast and add your sources — a YouTube URL, a PDF, a web link, pasted text, or several of them combined.
  2. Open Advanced Settings.
  3. Under Podcast Style, choose References & Key Data Recall.
  4. Pick a duration. Recaps compress well; a two-hour source rarely needs more than a short episode to list what it named.
  5. Generate.

Plan requirement, plainly: this style needs Creator or Pro. The Free tier cannot select it — the option is locked rather than quietly substituted, so you will not discover after the fact that you received a different format. Plan details are on the pricing page.


Which sources it suits — and which it does not

Strong fits. Interview and panel podcasts, where recommendations arrive in passing. Conference talks and lectures, where the slide with the reading list is on screen for four seconds. Book-club and review videos. Long unstructured essays that mention twenty things and footnote none of them. Multi-source generations, where you point it at four episodes of the same show at once and get one consolidated list.

Weak fits. Anything that already ships a reference list. A journal article ends with its bibliography, so running this style over one mostly gives you back what the PDF already gave you — for papers, research papers to podcast is the more useful route. Narrative and news content is a weak fit too, simply because there is little to extract.


A workflow that holds up

Run the recap first, before the full episode. It takes a couple of minutes and tells you whether the source is dense enough to be worth two hours of your attention — a triage step that is worth more than the recap itself on a bad source.

Then let it accumulate. Point the style at a batch of episodes from a show you follow and you end up with a standing reading list built out of things people you trust actually recommended, rather than things an algorithm surfaced. Because it is audio, that list survives a commute, and re-listening to it a week later is spaced repetition at essentially no cost.

Keep one habit: when a lead matters, check it. The style tells you where to look. It does not tell you that the spelling is right.


Frequently asked questions

Which plans include References & Key Data Recall?
Creator and Pro. The Free tier cannot select this style — it is gated at the plan level, so the option is locked rather than silently swapped for another format. If you are on Free and want it, the upgrade path is on the pricing page.
How accurate are the book titles and author names it returns?
Treat them as leads, not citations. Spoken names are the hardest part of any transcription pipeline: the Earnings-21 benchmark was built specifically to measure this and found that ASR accuracy for certain named-entity categories is poor. Unfamiliar surnames and foreign-language titles are the most likely to come back mis-spelled. Always verify anything you intend to cite or buy.
Can I use it on a two-hour YouTube interview?
Yes, and that is the case it was built for. Paste the video URL as a source. The style is most useful exactly where manual note-taking fails — long, unscripted, multi-guest conversations where references are name-dropped in passing and never appear in the description.
How many voices does the style use?
It defaults to a single narrator. A reference recap is a list read aloud, not a conversation, so a second voice adds turn-taking overhead without adding information.
Does it work in languages other than English?
Yes. Podcast style and output language are independent settings in Podhoc, so you can run an English-language video through References & Key Data Recall and receive the recap in Spanish, German or Arabic. Bear in mind that book titles are often left in their original language, which is usually what you want when you are searching for them later.
Will the generated recap be published publicly?
By default, Podhoc auto-publishes generated podcasts to its public Discover feed. Because this style is Creator and Pro only, everyone who can use it can also switch auto-publish off — per podcast or globally in account settings. If your sources are private or client-confidential, set that before you generate, not after.