Example: barber

HEIDELTIME.HR: Extracting and Normalizing …

: Extracting and NormalizingTemporal Expressions in CroatianLuka Skukan, Goran Glava , Jan najderUniversity of Zagreb, Faculty of Electrical Engineering and ComputingText Analysis and Knowledge Engineering LabUnska 3, 10000 Zagreb, Croatia{ , , expression extraction and normalization are important for many NLP tasks and have been the topic of extensive research. Whilethe majority of research on temporal expression extraction was performed for English, there has recently also been work on temporalprocessing for other languages. In this paper, we describe , the croatian resources for HeidelTime a multilingual,cross-domain temporal expression tagger. HeidelTime recognizes temporal expressions in text and normalizes them according to theTIMEX3 annotation standard. We compile WikiWarsHr, a corpus of historical narratives in croatian manually annotated for temporalexpressions. On WikiWarsHr, results comparable to those originally achieved by HeidelTime on Englishtexts, with F1-scores of and for expression extraction and normalization, : lu cenje in normaliziranje casovnih izrazov v hrva ciniLu cenje in normalizacija casovnih izrazov sta pomembna za raznovrstne naloge s podro cja ra cunalni ke obravnave naravnega jezika insta bila predmet tevilnih raziskav.}

HEIDELTIME.HR: Extracting and Normalizing Temporal Expressions in Croatian Luka Skukan, Goran Glavaš, Jan Šnajder University of Zagreb, Faculty of Electrical Engineering and Computing

Tags:

  Expression, Temporal, Normalizing, Extracting, Croatian, Extracting and normalizing temporal expressions in croatian

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of HEIDELTIME.HR: Extracting and Normalizing …

1 : Extracting and NormalizingTemporal Expressions in CroatianLuka Skukan, Goran Glava , Jan najderUniversity of Zagreb, Faculty of Electrical Engineering and ComputingText Analysis and Knowledge Engineering LabUnska 3, 10000 Zagreb, Croatia{ , , expression extraction and normalization are important for many NLP tasks and have been the topic of extensive research. Whilethe majority of research on temporal expression extraction was performed for English, there has recently also been work on temporalprocessing for other languages. In this paper, we describe , the croatian resources for HeidelTime a multilingual,cross-domain temporal expression tagger. HeidelTime recognizes temporal expressions in text and normalizes them according to theTIMEX3 annotation standard. We compile WikiWarsHr, a corpus of historical narratives in croatian manually annotated for temporalexpressions. On WikiWarsHr, results comparable to those originally achieved by HeidelTime on Englishtexts, with F1-scores of and for expression extraction and normalization, : lu cenje in normaliziranje casovnih izrazov v hrva ciniLu cenje in normalizacija casovnih izrazov sta pomembna za raznovrstne naloge s podro cja ra cunalni ke obravnave naravnega jezika insta bila predmet tevilnih raziskav.}

2 Medtem ko je bila ve cina raziskav lu cenja casovnih izrazov opravljenih za angle cino, pa so bilev zadnjem casu raziskave izvedene tudi za druge jezike. V prispevku opi emo , hrva ke vire za HeidelTime ve cjezi cniin prekdomenski ozna cevalec za casovne izraze. HeidelTime prepozna casovne izraze v besedilu in jih normalizira glede na standard zaozna cevanje TIMEX3. Izdelamo WikiWarsHr, korpus zgodovinskih pripovedi v hrva cini, ki je bil ro cno ozna cen za casovne izraze. NaWikiWarsHr dose e rezultate, primerljive s tistimi, ki jih je HeidelTime dosegal na angle kih besedilih, z mero F 0,93 zalu cenje in 0,86 za normalizacijo casovnih IntroductionThe ability to extract and normalize temporal expres-sions in natural language texts is of major importance fornatural language processing tasks, such as summarizationand question answering, but also for reasoning about eventsand time in general. temporal expression extraction is thetask of identifying temporal expressions and their normalization task amounts to turning extracted tem-poral expressions into a fully specified value and formattingthem according to some standard, including a number of temporal taggers are available,mostly for English and other major languages, a temporalexpression tagger for croatian does not yet exist.

3 A newtemporal expression tagger could be implemented, or anexisting multilingual system could be adapted to work forCroatian. We chose the latter approach in this work, build-ing on an existing and widely used this paper, we describe , the Croa-tian resources for the rule-based temporal expression tag-ger HeidelTime (Str tgen et al., 2013).1 HeidelTime ex-tracts and normalizes temporal expressions according to theTIMEX3 standard (Pustejovsky et al., 2003), and emergedas a winner in the TempEval-2 (Verhagen et al., 2010)and TempEval-3 (UzZaman et al., 2012) shared evalua-tion tasks. HeidelTime is a multilingual tagger, with re-sources been developed for English, German (Str tgen etal., 2013), Arabic, Italian, Spanish, Vietnamese (Str tgen1 al., 2014a), French (Moriceau and Tannier, 2014), Chi-nese (Li et al., 2014), Dutch, and Russian. We have devel-oped croatian resources, which will be included in the nextHeidelTime develop and evaluate the tagger, we compiled Wiki-WarsHr, a corpus of historical narratives in croatian man-ually annotated for temporal expressions.

4 On this cor-pus, results comparable to thoseoriginally achieved by HeidelTime on English structure of this paper is as follows. We describethe mechanisms of HeidelTime in Section 2. Section 3 de-scribes the In Section 4, wedescribe the WikiWarsHr corpus and present the evaluationresults. Section 5 concludes the HeidelTime taggerThe HeidelTime tagger extracts and normalizes tempo-ral expressions according to the TIMEX3 standard (Puste-jovsky et al., 2003). In TIMEX3, each temporal expres-sions is assigned a Type and a Value. A Type may be aDate,Time,DurationorSet. The Value corresponds to atemporal value, partially dependent on Type ( a Date 2014-10 for October of 2014).HeidelTime features a generic, language-independentcore, written in Java, and a language-dependent part, theso-called language resources. A language resource con-sist of three sets: (1) expression resources , (2) normaliza-2 The are also available KONFERENCA JEZIKOVNE TEHNOLOGIJE Informacijska dru ba - IS 20149th Language Technologies Conference Information Society - IS 201499(DCT: June 21st 2014)The field of AI research was founded at a conferenceon the campus of Dartmouth College in the <TIMEX3tid= t1 type= DATE value= 1956-SU >summer of1956</TIMEX3>.

5 <TIMEX3 tid= t2 type= DATE value= 2014 >58 years later</TIMEX3>., we stillhaven t achieved many of the goals proposed ,artificial intelligence has advanced and is<TIMEX3 tid= t3 type= DATE value= 2014-06-21 >today</TIMEX3> a part of our daily lives withoutmost of us knowing 1: Example of under-specification resources, and (3) rule resources. expression resourcesare regular expressions used for extraction temporal expres-sions from text, , phrases for months, weekdays, num-bers, etc. Normalization resources translate matched tokensto their canonical form, according to TIMEX3, by applyingnormalization mapping to extracted patterns ( , May 05 ). Finally, the rule resources combine the previoustwo resources to extract and normalize temporal expres-sions. These may be complemented with additional regularexpressions to form more complex match-and-normalizerules, , for discarding parts of extracted expressions orfor adding a modifier ( early , middle , etc.)

6 Normalization is performed both on fully specified ex-pressions ( June 28, 1995 ) and relative temporal expres-sions ( tomorrow ). The latter are expressions that cannotbe normalized without contextual information. Normaliza-tion of relative temporal expressions is performed by leav-ing the expressions under-specified and relying on Heidel-Time s generic focus-tracking system to assign them a morespecific value. For example, given a document creationtime (DCT) of June 20th, 2014, the expression tomorrow might be resolved as 2014-06-21 . This step is performedby taking into account the type of the document (narrative,news, scientific, or colloquial) and the tenses of the verbsused in the sentence containing the under-specified tempo-ral expression . Either the DCT or a previously mentionedvalue can be used in under-specified expression normaliza-tion, depending on the document type and the normaliza-tion rule. An example of resolving under-specified datesusing both DCT and current focus is shown in Fig.

7 1. Ad-ditionally, HeidelTime supports functionality extensions inform of text post-processors written as Java code. Theseallow for more verbose expression resolution, , com-puting the date of lunar holidays such as task of developing resources for croatian languageconsisted of developing three above-mentioned sets of re-sources. We next describe the resources and the develop-ment PreprocessingHeidelTime requires text to be pre-annotated with to-ken, sentence and part-of-speech (POS) information. Weused the CSTL emma lemmatiser (Jongejan and Haltrup,2005) for token splitting and lemmatization,3and the Hun-Pos part-of-speech tagger (Hal csy et al., 2007) to ob-tain the POS information. To integrate this functionalitywith HeidelTime, we wrote a Java wrapper that allows thetagger s engine to invoke it during pre-processing. Hun-Pos and CSTL emma were previously trained to work withCroatian texts (Agi c et al.)

8 , 2013). are divided into severalclasses. The expression and normalization resources aredivided into descriptive classes, according to their com-mon roles in temporal constructs, with each normalizationresource corresponding to an expression resource. Someexamples include rules are divided according to their semantics in theTIMEX3 standard intoDate,Time,DurationandSetre-sources. Altogether, there are 199 rule resources for Croat-ian: 123 for dates, 37 for time, 24 for durations, and 15 forsets. This number is much larger than for English, but ofcomparable size to resources for other inflected languages,such as French, which has 157 rule resources (Moriceauand Tannier, 2014). Furthermore, as a highly inflected lan-guage, croatian requires a large number of rule variationsto account for the inflections. This issue could have beenpartially avoided by using lemmas instead of raw , we chose not to do so for three reasons: (1) Theimplementation would be complex and time-consuming;(2) Due to the generic nature of the HeidelTime engine,lemmatization would have to be integrated system-wide,and the decision of whether to use lemmatization wouldhave to be specified for each set of language-specific re-sources.

9 (3) Errors in lemmatization would propagate intoHeidelTime, decreasing its an illustration, consider the following example of acomplete HeidelTime extraction rule, which can be used toextract and normalize parts of seasons, such as ranog pro-lje ca( early spring ):RULENAME="date_r9b",EXTRACTION="%rePar tWordsg1%reSeasong2",NORM_VALUE="UNDEF-y ear-%normSeason(g2)",NORM_MOD="%normPart Words(g1)"The extraction part of the rule extracts expressions de-scribing a specific part of something ( , early , mid-dle of , etc.) and stores it asgroup 1(g1), as well asan expression denoting a season ( , summer ), whichis stored asgroup 2(g2). It leaves the year undefined as UNDEF-year , which will be resolved by HeidelTime us-ing the temporal context of the sentence. The part word ing1is normalized as part of the modifier (NORM_MOD),which makes the value more specific. The season,g2,is combined with the determined year to get the tempo-ral value of the expression .

10 Assuming the inferred yearis 2014, the given expression ranog prolje ca would3 The lemmas produced by the CSTL emma lemmatiser arepresently not used by the system, but may be integrated in thefuture (cf. Section ).9. KONFERENCA JEZIKOVNE TEHNOLOGIJE Informacijska dru ba - IS 20149th Language Technologies Conference Information Society - IS 2014100be normalized as<TIMEX3 tid= t1 value= 2014-SP mod= START >ranog prolje ca</TIMEX3>. Development methodologyWe developed two phases. Wefirst translated the existing English and German re-sources (Str tgen et al., 2013) into croatian , wherever ap-propriate. We then used a data-driven approach to fur-ther develope and refine the resources, using a subset ofmanually-annotated Wikipedia corpus (cf. Section ) asa development set. The development set consists of tenWikipedia articles of varying length, altogether containing29,563 non-punctuation tokens and 677 temporal expres-sions.


Related search queries