Improved utilization methodology of BERT specialized in text classification

  • So, H.
  • Rhee, J.
Citations

SCOPUS

1

초록

Recent language models are pre-trained to generate universal word representations. This study proposes a BERT-Triplet model and its utilization methodology to generate word representations specialized for the text classification task. Specifically, we use class information of the data in the pre-training stage of the proposed BERT-Triplet model to closely distribute the embedding vectors of words or sentences with a high probability of being classified into the same class in the vector space, unlike existing language models. The proposed methodology obtains improvement of the classification performance and is expected to be used in various sub-fields of text classification and in language models other than BERT. © 2021 ACM.

키워드

BERTText classificationTriplet lossWord embedding
제목
Improved utilization methodology of BERT specialized in text classification
저자
So, H.Rhee, J.
DOI
10.1145/3471985.3472384
발행일
2021-03-05
유형
Conference Paper
저널명
ACM International Conference Proceeding Series
페이지
114 ~ 118