북 말뭉치

BookCorpus

북코퍼스(Toronto Book Corpus라고도 함)는 인터넷에서 스크랩된 약 11,000권의 미출판 서적의 텍스트로 구성된 데이터 세트입니다.OpenAI[1]의해 초기 GPT 모델을 훈련시키는 데 사용된 주요 말뭉치였으며, Google[2]BERT를 포함한 다른 초기 대형 언어 모델의 훈련 데이터로 사용되었습니다. OpenAI OpenAI OpenAI데이터 세트는 약 9억 8천 5백만 개의 단어로 구성되어 있으며, 이를 구성하는 책들은 로맨스, 공상과학,[2] 판타지를 포함한 다양한 장르에 걸쳐 있습니다.

이 말뭉치는 토론토 대학과 MIT의 연구원들이 2015년에 발표한 "책과 영화의 정렬:영화를 보고 책을 읽음으로써 이야기와 같은 시각적 설명을 향해"저자들은 그것을 "아직 출판되지 않은 [3][4]작가들이 쓴 무료 책들"로 구성되어 있다고 설명했습니다.데이터 세트는 처음에 토론토 대학 [4]웹 페이지에서 호스팅되었습니다.BookCorpusOpen이라는 하나 이상의 대체물이 생성되었지만 원본 데이터 세트의 공식 버전은 더 [5]이상 공개되지 않습니다.원래 2015년 논문에는 기록되지 않았지만, 말뭉치의 책이 스크랩된 사이트는 현재 스매시워드로 [4][5]알려져 있습니다.

레퍼런스

  1. ^ "Improving Language Understanding by Generative Pre-Training" (PDF). Archived (PDF) from the original on January 26, 2021. Retrieved June 9, 2020.
  2. ^ a b Devlin, Jacob; Chang, Ming-Wei; Lee, Kenton; Toutanova, Kristina (11 October 2018). "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding". arXiv:1810.04805v2 [cs.CL].
  3. ^ Zhu, Yukun; Kiros, Ryan; Zemel, Rich; Salakhutdinov, Ruslan; Urtasun, Raquel; Torralba, Antonio; Fidler, Sanja (2015). Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books. Proceedings of the IEEE International Conference on Computer Vision (ICCV).
  4. ^ a b c Lea, Richard (28 September 2016). "Google swallows 11,000 novels to improve AI's conversation". The Guardian.
  5. ^ a b Bandy, John; Vincent, Nicholas (2021). "Addressing "Documentation Debt" in Machine Learning: A Retrospective Datasheet for BookCorpus" (PDF). Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks.