Teach Your Web Page to Speak
You need no previous experience with speech technology. We will begin with one tiny instruction, understand every word, and gradually build a reliable talking avatar using JavaScript and the browser's built-in Speech Synthesis system.
What is speech synthesis?
Speech synthesis means making a computer produce spoken words. When written text is converted into speech, we also call it Text to Speech (TTS).
Your browser may already contain a speech engine and several voices. JavaScript sends text and settings to that engine. The speakers or headphones play the result.
Goal: By the end, a visitor will press Start, hear several paragraphs, see the active paragraph, and watch an avatar indicate when it is speaking.
Speak one sentence
Open an HTML file, put the following button inside it, and add the script below the button.
<button id="speakBtn">Speak</button>
<script>
const button = document.getElementById("speakBtn");
button.addEventListener("click", function () {
const message = new SpeechSynthesisUtterance("Hello, Champak!");
window.speechSynthesis.speak(message);
});
</script>getElementByIdfinds the button whose ID isspeakBtn.addEventListenerwaits for a click.SpeechSynthesisUtterancecreates a speech message. An utterance is something that is spoken.speechSynthesis.speak()places the message in the browser's speaking queue.
Speak whatever the user writes
Instead of fixing the sentence inside JavaScript, read it from a text box.
<textarea id="speechText">Welcome to Programmer's Picnic.</textarea>
<button id="readBtn">Read aloud</button>
<script>
const textBox = document.getElementById("speechText");
const readButton = document.getElementById("readBtn");
readButton.addEventListener("click", function () {
const text = textBox.value.trim();
if (text === "") {
alert("Please write something first.");
return;
}
const message = new SpeechSynthesisUtterance(text);
window.speechSynthesis.speak(message);
});
</script>value obtains the text inside the box. trim() removes unused spaces from its beginning and end. return stops the function when the box is empty.
Language, rate, pitch and volume
A speech message is an object. An object can store several related properties.
const message = new SpeechSynthesisUtterance("Welcome to Varanasi.");
message.lang = "en-IN"; // Indian English
message.rate = 1; // Speaking speed
message.pitch = 1; // How high or low the voice sounds
message.volume = 1; // 0 is silent; 1 is full volume
window.speechSynthesis.speak(message);en-IN for Indian English or hi-IN for Hindi.Pause, resume and stop
// Pause the current speech
window.speechSynthesis.pause();
// Continue paused speech
window.speechSynthesis.resume();
// Remove all speech from the queue
window.speechSynthesis.cancel();Pause keeps the current position. Resume continues from that position. Cancel stops the current message and clears messages waiting in the queue.
Speech events
An event is something that happens. An event handler is a function that runs when it happens.
const message = new SpeechSynthesisUtterance("Events make speech interactive.");
message.onstart = function () {
statusBox.textContent = "Speaking...";
};
message.onend = function () {
statusBox.textContent = "Finished.";
};
message.onerror = function (event) {
statusBox.textContent = "Problem: " + event.error;
};
window.speechSynthesis.speak(message);onstartruns when speaking begins.onendruns after the whole message finishes.onerrorruns if speech cannot continue.onboundarymay run at word or sentence boundaries, although browser support varies.
Choose a voice
Different devices contain different voices. We must ask the browser which voices are available instead of assuming a particular voice exists.
const voiceBox = document.getElementById("voiceBox");
let voices = [];
function loadVoices() {
voices = window.speechSynthesis.getVoices();
voiceBox.innerHTML = "";
voices.forEach(function (voice, index) {
const option = document.createElement("option");
option.value = index;
option.textContent = voice.name + " (" + voice.lang + ")";
voiceBox.appendChild(option);
});
}
loadVoices();
window.speechSynthesis.onvoiceschanged = loadVoices;Some browsers load their voice list later. Therefore we call loadVoices() immediately and also connect it to onvoiceschanged.
Build the complete talking avatar
Now we combine messages, events, controls, paragraph highlighting and animation. The mouth is a speaking indicator; it does not reproduce real human lip movements.
Lesson paragraphs
How the reliable version works
- Start always begins at paragraph one.
- A session number identifies the newest command.
- Callbacks and timers from cancelled speech are ignored.
- Cancellation errors do not accidentally skip a paragraph.
- Next completes the lesson after the final paragraph.
- Status changes are announced to assistive technology.
- Animation respects the user's reduced-motion preference.
What can you build next?
- A Read Aloud button on every lesson.
- A bilingual English and Hindi story reader.
- A pronunciation practice page with adjustable speed.
- A quiz that reads questions and answers.
- An accessible news or article reader.
- A photo-based presenter. Remember that realistic lip synchronization needs audio analysis or phoneme timing beyond this simple API.