한국어 중의성 사전 구성을 위한 기초 자료 연구
김수정
서울대 어학연구소
국어교육학연구 10집 407-427 (2000)
초록
중의성(ambiguity )이란 자연언어(NLP )의 커다란 특징 중의 하나로, 하나의 단어(w ord)가 여러 의미(m eanin g )로 해석될 수 있음을 말한다. 자연어 처리(Natural Lan guag e Processing )에서 발생하는 중의성 (Ambiguity )은 크게 구조적 중의성(structural ambiguity )과 어휘적 중 의성(lexical ambiguity )으로 나뉜다. 한국어의 중의성은 조사와 어미(Affix )의 기능에 의해 문장 성분 (sentece compon ent )이 결정되고 불규칙 활용(irregular conju gation ) 및 축약 현상이 발생하는 첨가어적인 특성, 즉 조사, 어미의 발달에 기 인한다. 본고에서 해결하고자 하는 한국어 중의성은 어휘적 중의성 (lexical ambiguity )을 비롯하여, 조사·어미(이상 Affix )가 결합하여 미 세하게 변이된 어절 중의성 등을 포함한다. 특히 Afffix 결합 중의성은 1 불규칙 용언 중의성 2 축약형 중의성 3 신조어 중의성 4 접사 결 합형 중의성 등으로 세분화된다. 어절 중의성 해결을 첫 번째 절차는 품 사 정보를 이용한 것이다. 동사/ 형용사 중의성의 경우에는 서술어가 요 구하는 논항(Argum ent ), 즉 격 정보를 알 수 있으므로 쉽게 해결될 수 있다. 그 다음의 절차는 해당 중의어의 의미자질 부여와 전후 어휘 문맥 정보를 이용하는 방법이다. 특히 관형사 한, 두, 세, 네, 열 등은 모두 어절 중의어를 가진 특징이 있다. 이러한 중의어의 처리는 전후 어휘 문 맥 정보의 비중이 크게 차지하게 된다. 중의성 처리의 선결 과제는 해당 어절 중의어의 의미자질 부여와 전후 어휘 문맥 정보를 위한 의미자질 의 선정이라고 할 수 있다. 427 < A b s tra c t > B as ic Cont ent s f or Kore an A m big uity Lex icon K im , S o o Ju n g T his study focu ses on con structin g basic contents for dev elopin g Korean ambiguity lexicon . Korean ambiguity con sist s of structural ambiguity an d lexical ambiguity . Especially it is one of ch aracteristic solving problem s of NLP (Natural Langu age Processin g ). T his paper focu ses on lexical ambiguity includin g affix - bound ambiguity w ords. Korean ambiguity w ords are appeared by combin ed affix es (Josa an d Omi). Especially affix - boun d ambiguity w ords include 1 irregular - v erb ambiguity 2 contraction - type ambiguity 3 coin age ambiguity 4 suffix - bound ambiguity . Korean disambiguation of NLP hav e three steps ; Fir st , U sing part s of speech inform ation is av ailable for unit of senten ce con struction disambiguation . Second, it needs as signing sem antic features to it s ambiguity w ords . Fin ally , context inform ation s are av ailable for solution of ambiguity w ords concernin g NLP .
