Amazon Polly

AmazonText-to-SpeechMultilingualGenerally AvailableProprietaryvm-amz-002

About

Cloud text-to-speech service offering standard, neural, and generative engine tiers. Provides SSML control, lexicon management, speech marks for lip sync, and newscaster and conversational speaking styles.

Capabilities (5)

Standard / Neural / Generative engines

SSML support

Speech marks for lip sync

Newscaster style

Custom lexicons

161 chars

Speed1.0x

Pitch1.0

0:00.00

Key Highlights

Three engine tiers let you balance cost, quality, and expressiveness

Newscaster and conversational speaking styles for media and IVR

Speech marks output enables real-time lip sync in avatar applications

Use Cases

Audiobook Narration

Generate natural-sounding narration for long-form content with consistent voice quality.

Notification Systems

Deliver voice alerts and notifications with expressive, human-like speech synthesis.

Multilingual Content

Produce audio content in multiple languages from a single text source.

Real-Time Voice Chat

Power low-latency voice responses in interactive applications and games.

Code Example

// Amazon Polly — Text-to-Speech
import { synthesize } from "@arkitekton/voice";

const audio = await synthesize({
  model: "vm-amz-002",
  vendor: "amazon",
  input: "Hello, welcome to Arkitekton.",
  voice: "alloy",
  response_format: "mp3",
  speed: 1.0,
});

// Play the audio
const blob = new Blob([audio], { type: "audio/mp3" });
const url = URL.createObjectURL(blob);
const player = new Audio(url);
player.play();

Related Models

PersonaPlex 7B

NVIDIA

NeMo TTS

NVIDIA

Riva

NVIDIA

ACE (Avatar Cloud Engine)

NVIDIA

gpt-4o-realtime

OpenAI

gpt-4o-mini-realtime

OpenAI

Quick Stats

Languages33 supported

LicenseProprietary

Pricing$4.00 / 1M characters (Neural)

StatusGenerally Available

Vendor

Amazon

Cloud-native speech services within the AWS ecosystem

View all Amazon models

Documentation

View on Amazon Site

Audiobook Narration

Generate natural-sounding narration for long-form content with consistent voice quality.

Notification Systems

Deliver voice alerts and notifications with expressive, human-like speech synthesis.

Multilingual Content

Produce audio content in multiple languages from a single text source.

Real-Time Voice Chat

Power low-latency voice responses in interactive applications and games.

Code Example

// Amazon Polly — Text-to-Speech
import { synthesize } from "@arkitekton/voice";

const audio = await synthesize({
  model: "vm-amz-002",
  vendor: "amazon",
  input: "Hello, welcome to Arkitekton.",
  voice: "alloy",
  response_format: "mp3",
  speed: 1.0,
});

// Play the audio
const blob = new Blob([audio], { type: "audio/mp3" });
const url = URL.createObjectURL(blob);
const player = new Audio(url);
player.play();