10th edition of the Language Resources and Evaluation Conference (LREC), Portorož 2016

OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles

author: Pierre Lison, Norwegian Computing Center
published: July 28, 2016, recorded: May 2016, views: 1282

Report a problem or upload files

If you have found a problem with this lecture or would like to send us extra material, articles, exercises, etc., please use our ticket system to describe your request and upload the data.
Enter your e-mail into the 'Cc' field, and we will keep you updated with your request's status.

Lecture popularity: You need to login to cast your vote.

Description

We present a new major release of the OpenSubtitles collection of parallel corpora. The release is compiled from a large database of movie and TV subtitles and includes a total of 1689 bitexts spanning 2.6 billion sentences across 60 languages. The release also incorporates a number of enhancements in the preprocessing and alignment of the subtitles, such as the automatic correction of OCR errors and the use of meta-data to estimate the quality of each subtitle and score subtitle pairs.

Link this page

Would you like to put a link to this lecture on your homepage?
Go ahead! Copy the HTML snippet !

Write your own review or comment:

Comment:
Name:
Email address:
URL:

make sure you have javascript enabled or clear this field:

OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles

See Also:

Related content

Report a problem or upload files

Description

Link this page

Write your own review or comment: