2nd International Chinese Word Segmentation Bakeoff Data Release Release 1, 2005 11 18 Introduction This directory contains the training, test, and gold standard data used in the 2nd International Chinese Word Segmentation Bakeoff. Also included is the script used to score the results submitted by the bakeoff participants and the simple segmenter used to generate the baseline and topline data. File List gold/ Contains the gold standard segmentation of the test data along with the training data word lists. scripts/ Contains the scoring script and simple segmenter. testing/ Contains the unsegmented test data. training/ Contains the segmented training data. doc/ Contains the instructions used in the bakeoff. Encoding Issues Files with the extension ".utf8" are encoded in UTF 8 Unicode. Files with the extension ".txt" are encoded as follows: as Big Five (CP950) hk Big Five/HKSCS msr EUC CN (CP936) pku EUC CN (CP936) EUC CN is often called "GB" or "GB2312" encoding, though technically GB2312 is a character set, not a character encoding. Scoring The script 'score' is used to generate compare two segmentations. The script takes three arguments: 1. The training set word list 2. The gold st…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy