Skip to main content

Audio Alternative Formats for Higher Education: Turning Course Content into Listening

How university teams turn lecture material, reading lists and course packs into audio — alternative formats, UDL and what to ask any partner.

Why university teams are looking at audio

Universities already produce alternative formats. A disability service converts a chapter to structured text, a library sources an accessible edition, a module leader records a summary. The work is real, it is skilled, and it is usually reactive — triggered by a request, delivered against a deadline that the academic calendar set months earlier.

Audio is one of the formats in that family. This article is written for the people who would own a decision about it — learning technologists, digital education managers, librarians and open-education leads, and disability and assistive technology services — rather than for individual students. If you are a student or a researcher looking for the personal workflow, Podhoc for students and Podhoc for researchers cover that ground instead.


What are alternative formats in higher education?

An alternative format is a version of the same course material produced so that a learner who cannot use the standard version can still use it. Audio sits alongside Braille, large print, structured digital text and tagged PDF. The point is equivalence: the same content, a different way in.

This is not only good practice, it is a documented obligation in several jurisdictions. The EU’s Web Accessibility Directive requires public sector bodies to publish an accessibility statement and to operate a feedback mechanism that lets people “report accessibility issues or request information in alternative formats” (European Commission — Web Accessibility). The request path is assumed to exist; the question for an institution is how much skilled staff time it takes to service it well.


Does audio satisfy Universal Design for Learning?

More directly than most vendors bother to cite. CAST’s UDL Guidelines name this practice explicitly under Consideration 2.2, “Support decoding of text, mathematical notation, and symbols”, whose suggestions include “Allow the use of text-to-speech” and “Use digital text with an accompanying human voice recording (e.g., DAISY Talking Books)” (CAST, 2024 — Consideration 2.2). That is the framework endorsing an audio version of course text, not merely tolerating it.

The framing is not novel, either. A systematic review in Innovations in Education and Teaching International addresses podcasting in higher education specifically as a component of Universal Design for Learning (Gunderson & Cumming, 2023) — so a proposal that puts the two together is working inside an established line of enquiry rather than inventing one.

The wider principle sits under Guideline 1, Perception, which asks educators to ensure key information is “equally perceptible to all learners” by “offering the same information through different modalities (e.g., through vision, hearing, or touch)”, with Consideration 1.2 explicit: “Support multiple ways to perceive information” (CAST, 2024 — Guideline 1: Perception). One citation note for anyone writing this into policy: the current reference is CAST (2024), CAST Universal Design for Learning Guidelines version 3.0, released on 30 July 2024. The older checkpoint numbering that many institutional accessibility policies still quote is obsolete.

Audio is one modality among several, not a universal solution, and any vendor who tells you otherwise is selling. Audio on its own excludes deaf and hard-of-hearing learners, and it removes the ability to skim, re-read and search that many learners depend on. The defensible position is that audio ships alongside the text, never instead of it.

One compliance detail is worth knowing, because it removes an objection that often stalls these conversations early. WCAG 2.2 Success Criterion 1.2.1 applies “except when the audio or video is a media alternative for text and is clearly labeled as such”, where a media alternative for text is “media that presents no more information than is already presented in text” (W3C — Audio-only and Video-only (Prerecorded)). A clearly labelled audio version of a course pack that adds nothing beyond its source text does not itself need a separate transcript — the text it was generated from is the transcript. That exception says nothing about the source document’s own accessibility obligations, which stand either way.


Which rules require accessible course materials?

Three instruments come up most often in procurement conversations. They differ in who they bind and what they demand, so it is worth being precise about each.

United States: ADA Title II and the 2027 deadline

The US Department of Justice’s Title II web accessibility rule sets WCAG 2.1 Level AA as “the technical standard for state and local governments’ web content and mobile apps.” On 20 April 2026 the Federal Register published the Department’s Interim Final Rule extending the compliance date for state and local government entities with a total population of 50,000 or more to 26 April 2027, with smaller public entities and special districts moving to 26 April 2028 (ADA.gov — Title II Web Rule). Title II reaches “all State and local governments and all departments, agencies, special purpose districts, and other instrumentalities of State or local government” (ADA.gov — Title II Primer), the category a public university sits in. The Interim Final Rule moved the date, not the obligation.

European Union: the Web Accessibility Directive

Directive (EU) 2016/2102 applies to public sector bodies, which includes public universities. The harmonised technical standard is EN 301 549 v3.2.1 (European Commission — Web Accessibility), which includes “WCAG 2.1 Level AA verbatim without modifications for Web content” (W3C — Web Accessibility Laws and Policies: European Union). Member States had to transpose the directive by 23 September 2018. Alongside the accessibility statement, the directive requires that feedback mechanism through which alternative formats can be requested.

United Kingdom: the 2018 regulations and the revision trigger

The Public Sector Bodies (Websites and Mobile Applications) Accessibility Regulations 2018 came into force on 23 September 2018, and current GOV.UK guidance directs public sector bodies to WCAG 2.2 Level AA. Confirm which version your own institution is actually working to before scoping anything: the harmonised standard EN 301 549 incorporates WCAG 2.1 Level AA, and a good deal of sector guidance and internal policy still references 2.1. The distinction rarely changes what you do about audio, but it changes which checklist an audit runs against. The exemptions are the part worth reading closely: documents published before 23 September 2018 are exempt “unless users need them to use a service”, and intranet or extranet content published before 23 September 2019 is exempt until it is subject to a major revision (GOV.UK — Accessibility requirements for public sector websites and apps). A module rewrite is a major revision. Refreshing a course is therefore the moment its legacy exemption lapses — which makes course-update cycles the natural place to introduce a new format.


Where does audio genuinely fit the course-content pipeline?

Not everywhere. Audio earns its place where the material is linear, explanatory and already written to be read end-to-end, and it struggles where the material is a reference the learner navigates rather than consumes.

Three points in a typical pipeline fit well:

  • Lecture material written as prose — module notes, lecture summaries and narrative handouts convert cleanly, because they already have an argument that runs start to finish.
  • Reading lists — particularly open-access articles, where an audio orientation to a paper helps a student decide how to spend their reading time. This is the same workflow described in Podhoc for researchers.
  • Course packs and open educational resources — self-contained explanatory material that is already cleared for reuse is the lowest-friction starting point for a pilot, because the licensing question is answered before you begin.

Three that fit badly: problem sets and worked mathematics, where notation carries the meaning; anything the learner is expected to navigate non-linearly, such as a statutes handbook or a lab manual; and material whose licence does not permit processing by a third party. That last one is usually the blocker in practice, and it is worth resolving before any technical question.


How is generated audio different from text-to-speech?

Text-to-speech reads a document aloud in sequence, preserving its structure. Generated audio restructures the material into a spoken explanation — in Podhoc’s case, a discussion between voices in one of several pedagogical formats such as Didactic, Feynman Technique or Critique (the audio styles).

These are different tools for different jobs, and an institution generally needs both. Text-to-speech is the right answer when a learner needs faithful access to the exact document. Generated audio is the right answer when the goal is comprehension of the ideas — an orientation to a paper before reading it, or a revision pass after. The evidence on why the second format holds attention differently is covered in the science of audio learning.

Neither removes the obligation to produce an accessible source document. Generated audio is an addition to the format range, not a repair for an inaccessible original.


What can Podhoc actually do today?

Being precise here matters more than being impressive, so this section is deliberately narrow.

Content sources. The Podhoc API accepts publicly accessible URLs. It does not ingest files or raw text, and it cannot reach material behind an authenticated virtual learning environment. For an institution this is the single most important constraint to design around: a pilot starts with content that is already openly addressable — open educational resources, open-access reading lists, public course pages — or it does not start through the API at all. Integration patterns covers what that looks like in a workflow.

Languages. Podhoc lists 73 languages, and the output language is independent of the source language, so material written in one language can be generated as audio in another. Coverage maturity varies across that catalogue, so test the specific languages a programme teaches in rather than assuming parity.

Formats. Eight pedagogical audio styles, of which Didactic and Feynman Technique are the ones most institutional pilots reach for.

Where episodes go. Podhoc publishes generated episodes to its public Discover catalogue by default, and whether that default can be switched off depends on the account. For an institution this is a governance question, not a preference — establish the visibility setting that applies to your account before any course material goes through the platform.

What does not exist today. There is no LTI integration, no VLE or LMS connector, and no single-sign-on bridge into Moodle, Canvas, Blackboard or Brightspace. If a proposal depends on one of those, it depends on something that has not been built.


What should you ask any audio partner?

This list applies to any supplier in this space, Podhoc included. It is deliberately not a feature comparison.

  • Licensing. Does your licence for the source material permit processing it through a third-party service? For reading-list content this is usually a publisher question, and it is usually the one that decides the pilot’s scope.
  • Data handling. Where is the content processed and stored, for how long, and what is the deletion path? Get this in writing before content moves.
  • Default visibility. Where does output land by default, who can see it, and can that default be changed for your account?
  • Accompanying text. Does every audio output ship with the text it was generated from, so the format is additive rather than substitutive?
  • Language coverage. Which languages are production-ready as opposed to listed, and how would you verify that for the languages you teach in?
  • Human review. Who checks generated audio for accuracy before it reaches students, and how is that step resourced? Generated material needs review; a partner who implies otherwise is describing a risk, not a feature.

How would a pilot actually work?

Narrowly, and with material whose licensing is already settled. A workable shape is one school or department, a small number of modules, audio published alongside the existing text rather than replacing it, and a defined review step before anything reaches students. Run it across a single teaching block so there is a beginning and an end, and decide in advance what would count as a result — whether that is displaced manual conversion work, learner uptake, or simply whether staff would keep using it.

Start with open educational resources or open-access reading if you want the fastest route to something real, because it removes the licensing question from the critical path. The two related pieces on audio for neurodivergent students and audio for training providers cover adjacent versions of the same question.

If that is worth a conversation, we would rather run one properly than describe one.

Request a pilot for your institution, company or employer →


Frequently asked questions

What is an alternative format in higher education?
An alternative format is a version of the same course material produced so that a learner who cannot use the standard version can still use it. Audio sits alongside Braille, large print, structured digital text and tagged PDF. The EU Web Accessibility Directive requires public sector bodies to run a feedback mechanism through which people can request information in an accessible alternative format.
Does Universal Design for Learning actually endorse audio versions of course text?
Yes, explicitly. CAST Universal Design for Learning Guidelines version 3.0 (2024) list under Consideration 2.2 the suggestions “Allow the use of text-to-speech” and “Use digital text with an accompanying human voice recording (e.g., DAISY Talking Books)”. Consideration 1.2, “Support multiple ways to perceive information”, frames audio as an additional modality rather than a substitute.
Does an audio version replace the written course material?
No. Audio is an additional way to perceive the same content, not a substitute for it. Audio on its own excludes deaf and hard-of-hearing learners, and it removes the ability to skim, re-read and search, so it should always ship alongside the text rather than instead of it.
Does an audio version of a course pack need its own transcript?
Not necessarily. WCAG 2.2 Success Criterion 1.2.1 applies except when the audio is a media alternative for text and is clearly labelled as such, where a media alternative for text presents no more information than is already in the text. A clearly labelled audio version that adds nothing beyond its source text is covered by that text. The source document keeps its own accessibility obligations regardless.
How is generated audio different from text-to-speech?
Text-to-speech reads a document aloud in sequence. Podhoc generates a structured spoken discussion between voices that explains the material, which is a different listening experience and a different production step. Both are useful; they solve different problems and neither removes the need for an accessible source document.
Which team at a university usually owns this?
In practice it is shared between learning technology or digital education, the library and its open-education leads, and the disability or assistive technology service. Procurement and IT security join once a supplier is in scope. A pilot is easiest to run when one of those teams sponsors it and the others review it.
Can Podhoc take material from behind our VLE login?
Not through the Podhoc API. The public API accepts publicly accessible URLs. Anything behind an authenticated virtual learning environment is out of scope for that route, so a pilot should start with material that is already openly published, such as open educational resources, reading lists of open-access articles, or public course descriptions.
Where do generated episodes end up by default?
Podhoc publishes generated episodes to its public Discover catalogue by default, and whether that default can be switched off depends on the account. Confirm the visibility setting that applies to your account before putting any course material through the platform — for an institution this is the first thing a pilot should establish, not the last.
Which languages can the audio be generated in?
Podhoc lists 73 languages, and the output language is independent of the source language, so material written in one language can be generated as audio in another. Coverage maturity varies by language, so a pilot should test the specific languages a programme actually teaches in rather than assuming parity across the catalogue.