상세 보기
초록
Recent language models are pre-trained to generate universal word representations. This study proposes a BERT-Triplet model and its utilization methodology to generate word representations specialized for the text classification task. Specifically, we use class information of the data in the pre-training stage of the proposed BERT-Triplet model to closely distribute the embedding vectors of words or sentences with a high probability of being classified into the same class in the vector space, unlike existing language models. The proposed methodology obtains improvement of the classification performance and is expected to be used in various sub-fields of text classification and in language models other than BERT. © 2021 ACM.
키워드
- 제목
- Improved utilization methodology of BERT specialized in text classification
- 저자
- So, H.; Rhee, J.
- 발행일
- 2021-03-05
- 유형
- Conference Paper
- 저널명
- ACM International Conference Proceeding Series
- 페이지
- 114 ~ 118