VTranslate
Corpus attribution

With thanks to the people who share their work.

We are delighted to use these films and recordings in our test corpus.

Their creators and communities make it possible to test translation on real dialogue, across languages, writing systems, and subtitle timings. Thank you for sharing your work.

Sintel

Sintel — © Blender Foundation and the Sintel team. Film licensed under CC BY 3.0.

A 35-second excerpt of the opening dialogue, plus multilingual subtitles. Our corpus contains excerpts and derived test material. Film source on Internet Archive.

Subtitle text: Wikimedia Commons TimedText contributors, under CC BY-SA 3.0. Film and subtitle source on Wikimedia Commons.

Big Buck Bunny

Big Buck Bunny — © Blender Foundation and the Big Buck Bunny team. Film licensed under CC BY 3.0.

A 30-second excerpt, plus multilingual subtitles, for sparse-dialogue cases. Our corpus contains excerpts and derived test material. Film source on Internet Archive.

Subtitle text: Wikimedia Commons TimedText contributors, under CC BY-SA 3.0. Film and subtitle source on Wikimedia Commons.

Tatoeba contributors

Our speech corpus uses eight short recordings from Tatoeba. We also concatenate these utterances into a multilingual test video, and use the English recording in a generated scene with text and a tone.

Sentence text is shared under CC BY 2.0 FR. Audio licenses and recording credits are specific to each contributor; follow the individual source pages for those details.

Generated fixtures

Our generated visual scenes, displayed text, and synthetic tone are dedicated under CC0 1.0. Recordings used within them retain their original licenses and credits.