-
147
pages
-
English
-
Documents
Description
Statistical Machine Translation: Trends & Challengesnd2 International Conference on Arabic Language Resources & Toolsst21 April 2009Dr. Hany HassanProf. Andy WayIBM Cairo HLT GroupNCLT/CNGL, IBM EgyptSchool of Computing,Dublin City University,Dublin 9, Irelandhanyh@eg.ibm.comaway@computing.dcu.ieOverview: Part 1 (AW)14:15 – 16:15• Why Corpus-Based MT?• Corpora, and Matters Arising•L a n g u a g e M o d e lin g• Translation Models• Word and Phrase Alignments•Decoding•EvaluationOverview: Part 2 (HH)16:30 – 18:30• Factored Models• Discriminative Training• Supertag Models of SMT• Open-Source ToolsWhy Corpus-Based MT?• the (relative) failure of rule-based approaches• the increasing availability of machine-readable text• the increase in capability of hardware (CPU, memory, disk space) with decrease in cost Sine qua nonA prerequisite for Data-Driven MT (and also TM, which is not MT, but rather CAT):• Example-Based MT (EBMT) • Statistical MT (SMT) • Hybrid Models which use some probabilistic processingis a parallel corpus (or bitext) of aligned sentences. Corpus-Based MT is here to stayThese approaches are now mainstream:• More researchers are developing corpus-based systems;• 1st company to use SMT now exists: www.languageweaver.com;• Irish MT company Traslán (www.traslan.ie) uses EBMT;• In recent large-scale evaluations, corpus-based MT systems come first.Two caveats:• Most industrial systems are still rule-based (but cf ...
-
Publié par
-
Langue
English