바이오케이션
Biocuration바이오케이션(Biocation)은 생물의학 데이터, 정보 및 지식을 스프레드시트, 표 및 지식 [1][2]그래프와 같은 구조화된 형식으로 정리하는 데 전념하는 생명과학 분야이다.생물의학 지식의 생물화는 생물정보화자, 소프트웨어 개발자 및 생물정보화자의 협력에 의해 가능하며 생물데이터베이스 작업의 [1]기초가 된다.
직업으로서의 생물 활동
바이오쿨레이터는 생물 및 모델 유기체 데이터베이스에 [3][4]의해 전파되는 정보를 큐레이션, 수집, 주석 달기 및 검증하는 전문 과학자입니다.이것은 새로운 직업으로, 면역 에피토프 데이터베이스와 분석 [5][6]자원과 같은 데이터베이스의 작업 맥락에서 2006년 과학 문헌 연대에 처음 언급되었다.바이오쿨레이터는 일반적으로 (온톨로지 [7]등을 통해) 습식 연구실에서의 경험과 지식의 컴퓨터 표현이 혼합된 박사급이다.
바이오쿨레이터의 역할은 출판을 목적으로 하는 일차 생물학적 연구 데이터의 품질 관리, 원본 과학 문헌에서 데이터를 추출 및 정리, 강력한 쿼리와 생물학적 데이터베이스 상호 운용을 가능하게 하는 표준 주석 프로토콜과 어휘로 데이터를 기술하는 것을 포함한다.바이오쿨레이터는 큐레이션된 정보의 정확성을 보장하고 [6]연구소와 데이터 교환을 촉진하기 위해 연구원과 소통합니다.
바이오쿨레이터는 다양한 연구 환경에 존재하지만, 바이오쿨레이터로 자칭되지는 않을 수 있습니다.ELIXIR(European Life-Sciences Infrastructure for Biological Informatics Learning, Education and Training)[8] 및 GOBLET(Global Organization for Biobytics Learning, Education and Training)와 같은 프로젝트는 훈련을 촉진하고 생물을 [9][10]진로로서 지원한다.
2011년에 생물정보화는 이미 전문직으로 인정받았지만,[11] 생물데이터 큐레이터를 목표로 준비하는 정식 학위 과정은 없었다.이 분야의 성장과 함께, 케임브리지 대학과 EMBL-EBI는 공동으로 바이오케이션 [12]대학원 수료증을 제공하기 시작했습니다.이 인증은 바이오케이션을 [13]그 자체로 인정하는 단계로 간주되고 있습니다.바이오포케이션 수요가 증가하고 있으며 [14]대학원 프로그램에 의한 추가적인 바이오포케이션 교육이 필요하다.
ClinGen(ClinGen)과 같은 바이오 큐레이터를 채용하고 있는 조직에서는, 바이오 [15]코메이션에 특화된 재료와 트레이닝을 제공하는 경우가 많습니다.
생물학적 지식 기반
바이오코쿠레이터의 역할은 생물학적 지식 기반 분야에서 가장 잘 알려져 있다.UniProt 및 PDB와[17] 같은 데이터베이스는[16] 정보를 정리하기 위해 전문 바이오쿨레이터에 의존합니다.무엇보다도, 바이오쿨레이터는 중복된 [18]엔트리를 병합하는 등 데이터 품질을 개선하기 위해 노력하고 있습니다.
이러한 지식 기반 중 중요한 부분은 모델 유기체 데이터베이스로, 특정 종류의 유기체에 대한 정보를 큐레이션하기 위해 바이오코레이터에 의존합니다.모델 유기체 데이터의 주목할 만한 예로는 FlyBase,[19][20] PomBase 및 [21]ZFIN이 있으며, 이는 각각 Drosophila, Shychoscaromyces, Zebrafish에 대한 정보를 큐레이션하기 위한 것이다.
큐레이션 및 주석
바이오케이션은 의미론적으로 표준화된 방식으로 온라인 데이터베이스에 생물 정보를 통합하는 것으로, 적절한 고유 추적 가능 식별자를 사용하고 소스 및 출처를 포함한 필요한 메타데이터를 제공합니다.
온톨로지, 제어된 어휘 및 표준 이름
바이오쿨레이터는 일반적으로 공유 바이오메디컬 온톨로지(오픈 바이오메디컬 온톨로지 등 많은 생물학적 및 의학 지식 영역을 포괄하는 구조화되고 통제된 용어)의 작성과 개발에 참여하고 있습니다.이러한 영역은 유전체학 및 단백질학, 해부학, 동물 및 식물의 발달, 생화학, 대사 경로, 분류학적 분류, 돌연변이 표현형을 포함한다.기존 온톨로지의 다양성을 고려할 때, 적합한 [22]온톨로지를 선택하는 방법에 대한 연구자들의 방향을 정하는 가이드라인이 있다.
Unified Medical Language System은 생명과학 [23]분야에서 사용되는 수백만 개의 용어를 통합하고 배포하는 시스템 중 하나입니다.
바이오쿨레이터는 유전자 명명 가이드라인을 일관되게 사용하고 다양한 모델 유기체의 유전자 명명 위원회에 참여하며, 종종 HUGO 유전자 명명 위원회(HGNC)와 협력한다.그들은 또한 국제생화학분자생물학연합(IUBMB)의 명명위원회가 제공한 것과 같은 다른 명명 가이드라인을 시행하고 있는데, 그 중 한 예가 효소위원회 EC 번호이다.
보다 일반적으로, 영속적 식별자의 사용은 지역사회에 의해 칭찬되며, 따라서 명확성을 향상시키고 지식을 용이하게 한다.
DNA주석
예를 들어 게놈 주석에서는 존재론자와 컨소시엄에 의해 정의된 식별자가 게놈의 일부를 기술하기 위해 사용된다.예를 들어, 유전자 온톨로지(GO)는 우리가 특정 유전자에 대해 알고 있는 것을 설명하기 위해 사용되는 생물학적 과정을 위한 용어들을 큐레이션한다.
텍스트 주석
2021년 현재 생명과학의 의사소통은 여전히 영어나 독일어와 같은 자유로운 자연언어를 통해 이루어지며, 어느 정도 모호함을 가지고 지식을 연결하기 어렵게 만든다.그래서 생물학적 배열에 주석을 다는 것 외에도, 바이오쿨레이터들은 단어들을 고유 식별자에 연결하면서 텍스트에 주석을 달기도 합니다.이를 통해 의미를 명확히 하고 텍스트를 컴퓨터로 처리할 수 있게 됩니다.텍스트 주석의 한 가지 적용은 과학자가 [25]언급하고 있는 정확한 유전자를 특정하는 것이다.
공개적으로 사용할 수 있는 텍스트 주석을 통해 생물학자들은 생물의학 텍스트를 더욱 활용할 수 있다.유럽 PMC에는 다양한 소스의 텍스트 주석을 중앙 집중화하여 SciLite라는 그래픽 [26]사용자 인터페이스에서 사용할 수 있는 애플리케이션 프로그래밍 인터페이스가 있습니다.PubTator Central은 주석도 제공하지만 완전히 컴퓨터화된 텍스트 마이닝을 기반으로 하며 사용자 [27]인터페이스를 제공하지 않습니다.또한 ezTag [28]시스템과 같이 사용자가 관심 있는 생물의학 텍스트에 수동으로 주석을 달 수 있는 프로그램도 있습니다.
국제 생물화 협회(ISB(International Society for Biocation(ISB)
ISB(International Society for Biocation)는 비영리 단체로 "생물학 분야를 촉진하고 회의와 워크숍을 통해 정보를 교환할 수 있는 장을 제공한다."국제생물학회의(International Biocation Conference)에서 성장하여 2009년 [4]초에 설립되었습니다.
ISB는 지역 내 바이오코쿠레이터에게 바이오코쿠레이터 커리어 어워드(연간 수여)와 바이오코쿠레이터 커리어 어워드(연간 수여)를 수여하고 있습니다.
ISB의 공식 저널 데이터베이스(Database)는 데이터베이스와 생물 [29]구성에 관한 기사를 전문으로 다루는 장소입니다.
커뮤니티 큐레이션
기존에는 데이터를 데이터베이스에 통합하는 전담 전문가가 바이오시케이션을 수행해 왔습니다.커뮤니티 큐레이션은 공개된 데이터에서 지식의 보급을 개선하고 생물 형성의 확장성을 개선하기 위한 비용 효율적인 방법을 제공하기 위한 유망한 접근법으로 등장했다.경우에 따라서는 커뮤니티의 도움이 행사 [30]중에 진행되는 큐레이션 작업에 도메인 전문가를 소개하는 잼버에 활용되는 반면, 다른 경우에는 전문가와 [31]비전문가의 비동기적 기여에 의존합니다.
생물학적 데이터베이스
몇몇 생물학적 데이터베이스는 유전자 식별자를 출판물 또는 자유 텍스트와 연관짓는 것에서부터 염기서열과 기능 데이터에 대한 보다 체계적이고 상세한 주석까지 기능 큐레이션 전략에 어느 정도 저자의 기여를 포함하며 전문 바이오 큐레이터와 동일한 표준에 따라 큐레이션을 출력한다.모델 생물 데이터베이스의 대부분의 커뮤니티 큐레이션은 큐레이션 대상 객체에 대한 정확한 식별자를 효과적으로 얻거나 상세 큐레이션을 위한 데이터 유형을 식별하기 위해 공개된 연구의 원본 작성자에 의한 주석(퍼스트 패스 주석)을 포함한다.예를 들어 다음과 같습니다.
- WormBase는 사용자로부터 퍼스트패스 주석을 성공적으로 요청하고 마이크로퍼블리싱 [33]프로세스와 함께 저자 큐레이션을 통합했습니다.또한 WormBase는 텍스트 마이닝을 플랫폼에 통합하여 커뮤니티 [32]큐레이터에게 제안합니다.
- FlyBase는 새로운 [34]출판물의 저자에게 이메일을 보내 온라인 툴을 통해 기술된 유전자와 데이터 유형을 나열하도록 초대하고 유전자 요약 [35]단락을 작성하기 위해 커뮤니티를 동원했다.
PomBase와 같은 다른 데이터베이스는 출판물을 위해 매우 상세한 온톨로지 기반 주석을 제출하고 통제된 어휘를 사용하여 게놈 전체 데이터 세트와 관련된 메타 데이터를 제출하기 위해 출판물 작성자에게 의존합니다.웹 기반 도구 Canto;[36]는 커뮤니티 제출을 용이하게 하기 위해 개발되었습니다.Canto는 자유롭게 사용할 수 있고 일반적이며 구성이 용이하기 때문에 다른 프로젝트에서 [37]채택되었습니다.큐레이션은 전문 큐레이터에 의해 검토되므로 모든 분자 데이터 [38]유형의 고품질 심층 큐레이션이 가능합니다.
널리 사용되는 UniProt 지식 기반은 또한 연구자들이 [39]단백질에 대한 정보를 추가할 수 있는 커뮤니티 큐레이션 메커니즘을 가지고 있다.
Wiki 스타일의 자원
바이오위키들은 콘텐츠를 제공하기 위해 커뮤니티에 의존하며, 생물 [40][41]서식화에 사용할 수 있는 일련의 위키 형식의 자원을 이용할 수 있습니다.예를 들어 AuthorReward는 [42]바이오위키에 대한 연구자들의 공헌을 수치화한 MediaWiki의 확장판이다.RiceWiki는 AuthorReward를 [43][44]탑재한 쌀 유전자 커뮤니티 큐레이션을 위한 위키 기반 데이터베이스의 한 예입니다.CAZypedia는 탄수화물 활성 효소(CAZys)[45]에 대한 정보의 공동 생물화를 위한 또 다른 위키이다.
WikiProteins/[46][47]WikiProfessional은 Barrend Mons가 주도하는 생물학적 데이터를 의미적으로 정리하는 프로젝트입니다.2007년 프로젝트는 Wikipedia 공동 설립자인 Jimmy Wales의 직접적인 공헌을 받아 Wikidata를 [46]영감으로 삼았다.현재 미디어위키 소프트웨어를 채택하여 실행되고 있는 프로젝트는 [48]생물학적 경로에 대한 정보를 크라우드소싱하는 WikiPathways입니다.
위키백과
과학 데이터베이스와 위키피디아 간의 경계가 점점 [49][41][50]모호해지면서 바이오코레이터와 위키피디아 사이에는 몇 가지 중복되는 부분이 있다.예를 들어 Rfam 및 Protein Data[53] Bank와 같은 데이터베이스는[51][52] 위키피디아와 편집자를 사용하여 정보를 [54][55]큐레이션합니다.그러나 대부분의 데이터베이스는 복잡한 조합으로 검색할 수 있는 고도로 구조화된 데이터를 제공하지만, Wikidata는 이 문제를 어느 정도 해결하는 것을 목표로 하고 있지만 일반적으로 Wikipedia에서는 가능하지 않습니다.
진 위키 프로젝트는 위키피디아를 수천 개의 유전자와 [56]티틴과 인슐린과 같은 유전자 산물의 공동 큐레이션에 이용했다.또한 여러 프로젝트에서는 의료 정보 [31]큐레이션을 위한 플랫폼으로 위키피디아를 사용한다.
Wikipedia가 생물 서식화에 사용되는 다른 한 가지 방법은 목록 기사를 통해서이다.예를 들어, 포괄적 항생제 내성 데이터베이스는 특정 위키피디아 [57]목록에 대한 항생제 내성에 대한 데이터베이스의 평가를 통합합니다.
위키데이터
Wikimedia Knowledge Base Wikidata는 생명과학 [58]전반에 걸친 통합 저장소로 생물정보화 커뮤니티에서 점점 더 많이 사용되고 있습니다.Wikidata는 소규모의 독립적인 생물학적 [59][60]지식 기반보다 유지관리 및 상호운용성에 대한 더 나은 전망을 가진 대안으로 일부에서는 보고 있다.
Wikidata는 SARS-CoV-2와 COVID-19[61][62] 대유행 및 유전자 [63]정보를 큐레이션하는 Gene Wiki 프로젝트에 의해 사용되어 왔다.Wikidata의 생물정보 데이터는 SPARQL [64]쿼리를 통해 외부 리소스에 재사용됩니다.일부 프로젝트에서는 Wikidata를 통한 큐레이션을 [65]Wikipedia의 생명과학 정보를 개선하기 위한 경로로 사용합니다.
게임화된 자원
게임 디자인 원리를 사용하여 참여도를 높이는 게이미티드 플랫폼을 통해 군중을 바이오 포지션에 참여시키는 접근법입니다.예를 들어 다음과 같습니다.
- Mark2Cure, 생물의학 추상화[66][67][68] 커뮤니티 큐레이션을 위한 게임화된 플랫폼
- 코크란 크라우드([69]Cochrane Crowd)는 임상시험 큐레이션과 생물의학 [70]문헌 분류 및 요약을 위한 코크란의 플랫폼입니다.
- CIViC는 점수를 추적하고[71] 순위표를 [72]유지하는 암과 관련된 게놈 변이체의 주석을 위한 포털이다.
- APICURON은 바이오코쿠레이터의 작업을 평가하고 인정하기 위한 데이터베이스로, 제3자 리소스로부터 생물포화 이벤트를 수집 및 집계하고 성과와 리더보드를 생성합니다.[73]
큐레이션을 위한 계산 텍스트 마이닝
자연어 처리 및 텍스트 마이닝 기술은 바이오쿨레이터가 수동 [75]큐레이션을 위해 정보를 추출하는 데 도움이 됩니다.텍스트 마이닝은 큐레이션 노력을 확장할 수 있으며, 예를 들어 유전자 이름의 식별을 지원하고 온톨로지를 [76][77]부분적으로 추론할 수 있다.구조화되지 않은 어설션을 구조화 정보로 변환하는 것은 명명된 엔티티 인식 및 의존관계 [78]해석과 같은 기술을 사용한다.생물의학 개념의 텍스트 마이닝은 리포트의 다양성에 관한 과제에 직면하고 있으며, 커뮤니티는 기사의 기계 [79]가독성을 높이기 위해 노력하고 있다.
COVID-19 대유행 기간 동안 생물의학 텍스트 마이닝은 이 주제에 대해 발표된 많은 양의 과학 연구를 처리하기 위해 많이 사용되었다(50,000개 이상의 기사).[80]
인기 있는 NLP 비단뱀 패키지 SpaCy는 Allen Institute for [81]AI에 의해 유지되는 생물의학 텍스트인 SciSpaCy를 수정했습니다.
바이오케이션에 적용되는 텍스트 마이닝의 과제 중 하나는 급여 장벽으로 인해 바이오메디컬 기사 전문에 접근하는 것의 어려움이며, 바이오케이션의 과제를 오픈 액세스 운동의 [82]과제와 연결시킨다.
텍스트 마이닝을 통한 생물 배치에 대한 보완적 접근법은 자동 주석 알고리즘과 결합된 생물의학 수치들에 광학 문자 인식을 적용하는 것을 포함한다.예를 들어,[83] 이것은 경로 수치에서 유전자 정보를 추출하기 위해 사용되어 왔다.
주석을 용이하게 하기 위해 쓰여진 텍스트를 개선하기 위한 제안은 통제된 자연[84] 언어를 사용하는 것에서부터 특정 관심 [84]종과 (유전자 및 단백질과 같은) 개념의 명확한 연관성을 제공하는 것까지 다양하다.
과제가 남아 있지만 텍스트 마이닝은 이미 여러 생물학적 지식 [85]기반에서 생물 배치 워크플로우의 필수적인 부분입니다.
생화학적인 과제
텍스트 마이닝과 바이오메이션 간의 인터페이스는 [86]2004년에 처음으로 발생한 일련의 텍스트 마이닝 경쟁인 BioCreAtIvE(생물학에서의 정보 추출 시스템의 중요 평가) 과제에 의해 촉진되었습니다.
「 」를 참조해 주세요.
레퍼런스
- ^ a b "What is biocuration? International Society for Biocuration". www.biocuration.org. Retrieved 2020-09-06.
- ^ Howe D, Costanzo M, Fey P, Gojobori T, Hannick L, Hide W, et al. (September 2008). "Big data: The future of biocuration". Nature. 455 (7209): 47–50. Bibcode:2008Natur.455...47H. doi:10.1038/455047a. PMC 2819144. PMID 18769432.
- ^ Burge S, Attwood TK, Bateman A, Berardini TZ, Cherry M, O'Donovan C, et al. (2012-03-20). "Biocurators and biocuration: surveying the 21st century challenges". Database. 2012: bar059. doi:10.1093/database/bar059. PMC 3308150. PMID 22434828.
- ^ a b Bateman A (April 2010). "Curators of the world unite: the International Society of Biocuration". Bioinformatics. 26 (8): 991. doi:10.1093/bioinformatics/btq101. PMID 20305270.
- ^ Bourne PE, McEntyre J (October 2006). "Biocurators: contributors to the world of science". PLOS Computational Biology. 2 (10): e142. Bibcode:2006PLSCB...2..142B. doi:10.1371/journal.pcbi.0020142. PMC 1626157. PMID 17411327.
- ^ a b Salimi N, Vita R (October 2006). "The biocurator: connecting and enhancing scientific data". PLOS Computational Biology. 2 (10): e125. Bibcode:2006PLSCB...2..125S. doi:10.1371/journal.pcbi.0020125. PMC 1626147. PMID 17069454.
- ^ Biocuration, International Society for (2018-04-16). "Biocuration: Distilling data into knowledge". PLOS Biology. 16 (4): e2002846. doi:10.1371/JOURNAL.PBIO.2002846. PMC 5919672. PMID 29659566.
- ^ "GOBLET The Global Organisation for Bioinformatics Learning, Education & Training". Retrieved 2020-12-19.
- ^ Alexandra Holinski; Melissa Burke; Sarah L Morgan; Peter McQuilton; Patricia M. Palagi (4 September 2020). "Biocuration - mapping resources and needs". F1000Research. 9: 1094. doi:10.12688/F1000RESEARCH.25413.1. ISSN 2046-1402. PMC 7590901. PMID 33145007. Wikidata Q101217428.
- ^ EMBL-EBI. "Biocuration EMBL-EBI Training". www.ebi.ac.uk. Retrieved 2022-05-06.
- ^ Sanderson, Katharine (February 2011). "Bioinformatics: Curation generation". Nature. 470 (7333): 295–296. doi:10.1038/nj7333-295a. ISSN 1476-4687. PMID 21348148.
- ^ Anonymous (2019-10-30). "Postgraduate Certificate in Biocuration". www.ice.cam.ac.uk. Retrieved 2020-10-06.
- ^ Tang YA, Pichler K, Füllgrabe A, Lomax J, Malone J, Munoz-Torres MC, et al. (May 2019). "Ten quick tips for biocuration". PLOS Computational Biology. 15 (5): e1006906. Bibcode:2019PLSCB..15E6906T. doi:10.1371/journal.pcbi.1006906. PMC 6497217. PMID 31048830.
- ^ Harper, Lisa; Campbell, Jacqueline D.; Cannon, Ethalinda K. S.; Jung, Sook; Poelchau, Monica F.; Walls, Ramona L.; Andorf, Carson M.; Arnaud, Elizabeth; Berardini, Tanya Z.; Birkett, Clayton; Cannon, Steve (2018-01-01). "AgBioData consortium recommendations for sustainable genomics and genetics databases for agriculture". Database. 2018: 1–32. doi:10.1093/DATABASE/BAY088. PMC 6146126. PMID 30239679.
- ^ "Biocurator - ClinGen Clinical Genome Resource". www.clinicalgenome.org. Retrieved 2021-05-26.
- ^ "UniProt: the universal protein knowledgebase". Nucleic Acids Research. 45 (D1): D158–D169. 2016-11-29. doi:10.1093/nar/gkw1099. ISSN 0305-1048. PMC 5210571. PMID 27899622.
- ^ Berman, Helen M.; Westbrook, J.; Feng, Z.; Gilliland, G.; Bhat, T. N.; Weissig, H.; Shindyalov, Ilya; Bourne, Philip (2000-01-01). "The Protein Data Bank". Nucleic Acids Research. 28 (1): 235–242. doi:10.1093/NAR/28.1.235. PMC 102472. PMID 10592235.
- ^ Chen, Qingyu; Britto, Ramona; Erill, Ivan; Jeffery, Constance J.; Liberzon, Arthur; Magrane, Michele; Onami, Jun-Ichi; Robinson-Rechavi, Marc; Sponarova, Jana; Zobel, Justin; Verspoor, Karin (2020-07-08). "Quality Matters: Biocuration Experts on the Impact of Duplication and Other Data Quality Issues in Biological Databases". Genomics Proteomics and Bioinformatics. 18 (2): 91–103. doi:10.1016/J.GPB.2018.11.006. PMC 7646089. PMID 32652120.
- ^ "FlyBase: a Drosophila database. Flybase Consortium". Nucleic Acids Research. 26 (1): 85–88. 1998-01-01. doi:10.1093/nar/26.1.85. ISSN 1362-4962. PMC 147222. PMID 9399806.
- ^ Lock, Antonia; Rutherford, Kim; Harris, Midori A; Hayles, Jacqueline; Oliver, Stephen G; Bähler, Jürg; Wood, Valerie (2018-10-13). "PomBase 2018: user-driven reimplementation of the fission yeast database provides rapid and intuitive access to diverse, interconnected information". Nucleic Acids Research. 47 (D1): D821–D827. doi:10.1093/nar/gky961. ISSN 0305-1048. PMC 6324063. PMID 30321395.
- ^ Ruzicka, Leyla; Howe, Douglas G.; Ramachandran, Sridhar; Toro, Sabrina; Slyke, Ceri E. Van; Bradford, Yvonne M.; Eagle, Anne; Fashena, David; Frazer, Ken; Kalita, Patrick; Mani, Prita (2019-01-01). "The Zebrafish Information Network: new support for non-coding genes, richer Gene Ontology annotations and the Alliance of Genome Resources". Nucleic Acids Research. 47 (D1): D867–D873. doi:10.1093/NAR/GKY1090. PMC 6323962. PMID 30407545.
- ^ Malone J, Stevens R, Jupp S, Hancocks T, Parkinson H, Brooksbank C (February 2016). "Ten Simple Rules for Selecting a Bio-ontology". PLOS Computational Biology. 12 (2): e1004743. Bibcode:2016PLSCB..12E4743M. doi:10.1371/journal.pcbi.1004743. PMC 4750991. PMID 26867217.
- ^ Bodenreider O (January 2004). "The Unified Medical Language System (UMLS): integrating biomedical terminology". Nucleic Acids Research. 32 (Database issue): D267-70. doi:10.1093/nar/gkh061. PMC 308795. PMID 14681409.
- ^ McMurry JA, Juty N, Blomberg N, Burdett T, Conlin T, Conte N, et al. (June 2017). "Identifiers for the 21st century: How to design, provision, and reuse persistent identifiers to maximize utility and impact of life science data". PLOS Biology. 15 (6): e2001414. doi:10.1371/journal.pbio.2001414. PMC 5490878. PMID 28662064.
- ^ Mons B (June 2005). "Which gene did you mean?". BMC Bioinformatics. 6 (1): 142. doi:10.1186/1471-2105-6-142. PMC 1173089. PMID 15941477.
- ^ Venkatesan A, Kim JH, Talo F, Ide-Smith M, Gobeill J, Carter J, et al. (2016-12-12). "SciLite: a platform for displaying text-mined annotations as a means to link research articles with biological data". Wellcome Open Research. 1: 25. doi:10.12688/wellcomeopenres.10210.1. PMC 5527546. PMID 28948232.
- ^ Wei CH, Allot A, Leaman R, Lu Z (July 2019). "PubTator central: automated concept annotation for biomedical full text articles". Nucleic Acids Research. 47 (W1): W587–W593. doi:10.1093/nar/gkz389. PMC 6602571. PMID 31114887.
- ^ Kwon D, Kim S, Wei CH, Leaman R, Lu Z (July 2018). "ezTag: tagging biomedical concepts via interactive learning". Nucleic Acids Research. 46 (W1): W523–W529. doi:10.1093/nar/gky428. PMC 6030907. PMID 29788413.
- ^ Landsman, D.; Gentleman, R.; Kelso, J.; Francis Ouellette, B. F. (2010-01-05). "DATABASE: A new forum for biological databases and curation". Database. 2009: bap002. doi:10.1093/database/bap002. ISSN 1758-0463. PMC 2790300. PMID 20157475.
- ^ Naithani, Sushma; Gupta, Parul; Preece, Justin; Garg, Priyanka; Fraser, Valerie; Padgitt-Cobb, Lillian K; Martin, Matthew; Vining, Kelly; Jaiswal, Pankaj (2019-01-01). "Involving community in genes and pathway curation". Database. 2019. doi:10.1093/database/bay146. ISSN 1758-0463. PMC 6334007. PMID 30649295.
- ^ a b Denise A. Smith (18 February 2020). Stefano Triberti (ed.). "Situating Wikipedia as a health information resource in various contexts: A scoping review". PLOS One. 15 (2): e0228786. doi:10.1371/JOURNAL.PONE.0228786. ISSN 1932-6203. PMC 7028268. PMID 32069322. Wikidata Q85632863.
- ^ a b Arnaboldi V, Raciti D, Van Auken K, Chan JN, Müller HM, Sternberg PW (January 2020). "Text mining meets community curation: a newly designed curation platform to improve author experience and participation at WormBase". Database. 2020. doi:10.1093/database/baaa006. PMC 7078066. PMID 32185395. S2CID 212750405.
- ^ Lee RY, Howe KL, Harris TW, Arnaboldi V, Cain S, Chan J, et al. (January 2018). "WormBase 2017: molting into a new stage". Nucleic Acids Research. 46 (D1): D869–D874. doi:10.1093/nar/gkx998. PMC 5753391. PMID 29069413.
- ^ Bunt SM, Grumbling GB, Field HI, Marygold SJ, Brown NH, Millburn GH (2012). "Directly e-mailing authors of newly published papers encourages community curation". Database. 2012: bas024. doi:10.1093/database/bas024. PMC 3342516. PMID 22554788.
- ^ Antonazzo G, Urbano JM, Marygold SJ, Millburn GH, Brown NH (January 2020). "Building a pipeline to solicit expert knowledge from the community to aid gene summary curation". Database. 2020. doi:10.1093/database/baz152. PMC 6971343. PMID 31960022.
- ^ Rutherford KM, Harris MA, Lock A, Oliver SG, Wood V (June 2014). "Canto: an online tool for community literature curation". Bioinformatics. 30 (12): 1791–2. doi:10.1093/bioinformatics/btu103. PMC 4058955. PMID 24574118.
- ^ "pombase/canto". PomBase. 25 September 2020.
- ^ Lock A, Harris MA, Rutherford K, Hayles J, Wood V (January 2020). "Community curation in PomBase: enabling fission yeast experts to provide detailed, standardized, sharable annotation from research publications". Database. 2020. doi:10.1093/database/baaa028. PMC 7192550. PMID 32353878.
- ^ "UniProt: the universal protein knowledgebase". Nucleic Acids Research. 45 (D1): D158–D169. 2016-11-29. doi:10.1093/nar/gkw1099. ISSN 0305-1048. PMC 5210571. PMID 27899622.
- ^ Khare, Ritu; Good, Benjamin M.; Leaman, Robert; Su, Andrew I.; Lu, Zhiyong (2016-01-01). "Crowdsourcing in biomedicine: challenges and opportunities". Briefings in Bioinformatics. 17 (1): 23–32. doi:10.1093/BIB/BBV021. PMC 4719068. PMID 25888696.
- ^ a b Finn RD, Gardner PP, Bateman A (January 2012). "Making your database available through Wikipedia: the pros and cons". Nucleic Acids Research. 40 (Database issue): D9-12. doi:10.1093/nar/gkr1195. PMC 3245093. PMID 22144683.
- ^ Dai L, Tian M, Wu J, Xiao J, Wang X, Townsend JP, Zhang Z (July 2013). "AuthorReward: increasing community curation in biological knowledge wikis through automated authorship quantification". Bioinformatics. 29 (14): 1837–9. doi:10.1093/bioinformatics/btt284. PMC 3702255. PMID 23732274.
- ^ Zhang Z, Sang J, Ma L, Wu G, Wu H, Huang D, et al. (January 2014). "RiceWiki: a wiki-based database for community curation of rice genes". Nucleic Acids Research. 42 (Database issue): D1222-8. doi:10.1093/nar/gkt926. PMC 3964990. PMID 24136999.
- ^ "Os01g0883800 - RiceWiki". 2017-10-20. Archived from the original on 2017-10-20. Retrieved 2020-09-06.
- ^ Consortium, CAZypedia (2017-10-11). "Ten years of CAZypedia: a living encyclopedia of carbohydrate-active enzymes". Glycobiology. 28 (1): 3–8. doi:10.1093/GLYCOB/CWX089. PMID 29040563.
- ^ a b Mons B, Ashburner M, Chichester C, van Mulligen E, Weeber M, den Dunnen J, et al. (2008-05-28). "Calling on a million minds for community annotation in WikiProteins". Genome Biology. 9 (5): R89. doi:10.1186/gb-2008-9-5-r89. PMC 2441475. PMID 18507872.
- ^ Giles J (February 2007). "Key biology databases go wiki". Nature. 445 (7129): 691. Bibcode:2007Natur.445..691G. doi:10.1038/445691a. PMID 17301755. S2CID 4410783.
- ^ "WikiPathways - WikiPathways". www.wikipathways.org. Retrieved 2020-10-14.
- ^ Wodak SJ, Mietchen D, Collings AM, Russell RB, Bourne PE (2012). "Topic pages: PLOS Computational Biology meets Wikipedia". PLOS Computational Biology. 8 (3): e1002446. Bibcode:2012PLSCB...8E2446W. doi:10.1371/journal.pcbi.1002446. PMC 3315447. PMID 22479174.
- ^ Page RD (March 2011). "Linking NCBI to Wikipedia: a wiki-based approach". PLOS Currents. 3: RRN1228. doi:10.1371/currents.RRN1228. PMC 3080707. PMID 21516242.
- ^ Gardner PP, Daub J, Tate J, Moore BL, Osuch IH, Griffiths-Jones S, et al. (January 2011). "Rfam: Wikipedia, clans and the "decimal" release". Nucleic Acids Research. 39 (Database issue): D141-5. doi:10.1093/nar/gkq1129. PMC 3013711. PMID 21062808.
- ^ Daub J, Gardner PP, Tate J, Ramsköld D, Manske M, Scott WG, et al. (December 2008). "The RNA WikiProject: community annotation of RNA families". RNA. 14 (12): 2462–4. doi:10.1261/rna.1200508. PMC 2590952. PMID 18945806.
- ^ Burkhardt K, Schneider B, Ory J (October 2006). "A biocurator perspective: annotation at the Research Collaboratory for Structural Bioinformatics Protein Data Bank". PLOS Computational Biology. 2 (10): e99. Bibcode:2006PLSCB...2...99B. doi:10.1371/journal.pcbi.0020099. PMC 1626146. PMID 17069453.
- ^ Logan DW, Sandal M, Gardner PP, Manske M, Bateman A (September 2010). "Ten simple rules for editing Wikipedia". PLOS Computational Biology. 6 (9): e1000941. Bibcode:2010PLSCB...6E0941L. doi:10.1371/journal.pcbi.1000941. PMC 2947980. PMID 20941386.
- ^ Butler D (2008). "Publish in Wikipedia or perish: Journal to require authors to post in the free online encyclopaedia". Nature. doi:10.1038/news.2008.1312.
- ^ Huss JW, Lindenbaum P, Martone M, Roberts D, Pizarro A, Valafar F, et al. (January 2010). "The Gene Wiki: community intelligence applied to human gene annotation". Nucleic Acids Research. 38 (Database issue): D633-9. doi:10.1093/nar/gkp760. PMC 2808918. PMID 19755503.
- ^ Alcock, Brian P.; Raphenya, Amogelang R.; Lau, Tammy T. Y.; Tsang, Kara K.; Bouchard, Mégane; Edalatmand, Arman; Huynh, William; Nguyen, Anna-Lisa V.; Cheng, Annie A.; Liu, Sihan; Min, Sally Y. (2020-01-01). "CARD 2020: antibiotic resistome surveillance with the comprehensive antibiotic resistance database". Nucleic Acids Research. 48 (D1): D517–D525. doi:10.1093/NAR/GKZ935. PMC 7145624. PMID 31665441.
- ^ Waagmeester A, Stupp G, Burgstaller-Muehlbacher S, Good BM, Griffith M, Griffith OL, et al. (March 2020). Rodgers P, Mungall C (eds.). "Wikidata as a knowledge graph for the life sciences". eLife. 9: e52614. doi:10.7554/eLife.52614. PMC 7077981. PMID 32180547. S2CID 212739087.
- ^ Rutz, Adriano; Sorokina, Maria; Galgonek, Jakub; Mietchen, Daniel; Willighagen, Egon; Gaudry, Arnaud; Graham, James G; Stephan, Ralf; Page, Roderic; Vondrášek, Jiří; Steinbeck, Christoph; Pauli, Guido F; Wolfender, Jean-Luc; Bisson, Jonathan; Allard, Pierre-Marie (26 May 2022). "The LOTUS initiative for open knowledge management in natural products research". eLife. 11: e70780. doi:10.7554/eLife.70780. PMID 35616633. S2CID 249064853.
- ^ Rutz, Adriano; Sorokina, Maria; Galgonek, Jakub; Mietchen, Daniel; Willighagen, Egon; Gaudry, Arnaud; Graham, James G.; Stephan, Ralf; Page, Roderic; Vondrášek, Jiří; Steinbeck, Christoph; Pauli, Guido F.; Wolfender, Jean-Luc; Bisson, Jonathan; Allard, Pierre-Marie (24 December 2021). "The LOTUS Initiative for Open Natural Products Research: Knowledge Management through Wikidata": 2021.02.28.433265. doi:10.1101/2021.02.28.433265. S2CID 235262250.
{{cite journal}}:Cite 저널 요구 사항journal=(도움말) - ^ Turki, Houcemeddine; Taieb, Mohamed Ali Hadj; Shafee, Thomas; Lubiana, Tiago; Jemielniak, Dariusz; Aouicha, Mohamed Ben; Gayo, José Emilio Labra; Youngstrom, Eric; Banat, Mossab; Das, Diptanshu; Mietchen, Daniel (2021-02-18). Haller, Armin (ed.). "Representing COVID-19 information in collaborative knowledge graphs: the case of Wikidata" (PDF).
- ^ Waagmeester, Andra; Willighagen, Egon L.; Su, Andrew I.; Kutmon, Martina; Gayo, Jose Emilio Labra; Fernández-Álvarez, Daniel; Groom, Quentin; Schaap, Peter J.; Verhagen, Lisa M.; Koehorst, Jasper J. (2021-01-22). "A protocol for adding knowledge to Wikidata: aligning resources on human coronaviruses". BMC Biology. 19 (1): 12. doi:10.1186/s12915-020-00940-y. ISSN 1741-7007. PMC 7820539. PMID 33482803.
- ^ Burgstaller-Muehlbacher S, Waagmeester A, Mitraka E, Turner J, Putman T, Leong J, et al. (2016). "Wikidata as a semantic framework for the Gene Wiki initiative". Database. 2016: baw015. doi:10.1093/database/baw015. PMC 4795929. PMID 26989148.
- ^ Willighagen, Egon; Martens, Marvin; Yasunori; Lubiana, Tiago; Nunogit; Mietchen, Daniel; Addshore (2020-08-09), egonw/SARS-CoV-2-Queries: Edition 1, doi:10.5281/zenodo.3977414, retrieved 2021-04-14
- ^ Alexander Pfundner; Tobias Schönberg; John Horn; Richard D Boyce; Matthias Samwald (5 May 2015). "Utilizing the Wikidata system to improve the quality of medical content in Wikipedia in diverse languages: a pilot study". Journal of Medical Internet Research. 17 (5): e110. doi:10.2196/JMIR.4163. ISSN 1438-8871. PMC 4468594. PMID 25944105. Wikidata Q21503276.
- ^ Tsueng G, Nanis SM, Fouquier J, Good BM, Su AI (2016-12-31). "Citizen Science for Mining the Biomedical Literature". Citizen Science. 1 (2): 14. doi:10.5334/cstp.56. PMC 6226017. PMID 30416754.
- ^ Tsueng G, Nanis M, Fouquier JT, Mayers M, Good BM, Su AI (February 2020). "Applying citizen science to gene, drug and disease relationship extraction from biomedical abstracts". Bioinformatics. 36 (4): 1226–1233. doi:10.1093/bioinformatics/btz678. PMC 8104067. PMID 31504205.
- ^ "Play Mark2Cure, help identify key terms in biomedical research abstracts". Citizen Science Games. Retrieved 2020-09-06.
- ^ "Cochrane Crowd". crowd.cochrane.org. Retrieved 2020-09-25.
- ^ Gartlehner G, Affengruber L, Titscher V, Noel-Storr A, Dooley G, Ballarini N, König F (May 2020). "Single-reviewer abstract screening missed 13 percent of relevant studies: a crowd-based, randomized controlled trial". Journal of Clinical Epidemiology. 121: 20–28. doi:10.1016/j.jclinepi.2020.01.005. PMID 31972274.
- ^ Griffith, Malachi; Spies, Nicholas C; Krysiak, Kilannin; McMichael, Joshua F; Coffman, Adam C; Danos, Arpad M; Ainscough, Benjamin J; Ramirez, Cody A; Rieke, Damian T; Kujan, Lynzey; Barnell, Erica K (2017-01-31). "CIViC is a community knowledgebase for expert crowdsourcing the clinical interpretation of variants in cancer". Nature Genetics. 49 (2): 170–174. doi:10.1038/ng.3774. hdl:10230/46299. ISSN 1061-4036. PMC 5367263. PMID 28138153.
- ^ "CIViC - Clinical Interpretation of Variants in Cancer". civicdb.org. Retrieved 2021-04-14.
- ^ Hatos, András; Quaglia, Federica; Piovesan, Damiano; Tosatto, Silvio C. E. (2021-04-21). "APICURON: a database to credit and acknowledge the work of biocurators". Database: The Journal of Biological Databases and Curation. 2021: baab019. doi:10.1093/database/baab019. ISSN 1758-0463. PMC 8060004. PMID 33882120.
- ^ Percha, Bethany; Altman, Russ B. (2018-08-01). "A global network of biomedical relationships derived from text". Bioinformatics. 34 (15): 2614–2624. doi:10.1093/bioinformatics/bty114. ISSN 1367-4803. PMC 6061699. PMID 29490008.
- ^ Hirschman L, Burns GA, Krallinger M, Arighi C, Cohen KB, Valencia A, et al. (2012). "Text mining for the biocuration workflow". Database. 2012: bas020. doi:10.1093/database/bas020. PMC 3328793. PMID 22513129.
- ^ Ananiadou, Sophia; Kell, Douglas B.; Tsujii, Jun-ichi (December 2006). "Text mining and its potential applications in systems biology". Trends in Biotechnology. 24 (12): 571–579. doi:10.1016/j.tibtech.2006.10.002. ISSN 0167-7799. PMID 17045684.
- ^ Winnenburg, R.; Wachter, T.; Plake, C.; Doms, A.; Schroeder, M. (2008-07-11). "Facts from text: can text mining help to scale-up high-quality manual curation of gene products with ontologies?". Briefings in Bioinformatics. 9 (6): 466–478. doi:10.1093/bib/bbn043. ISSN 1467-5463. PMID 19060303.
- ^ Percha, Bethany; Altman, Russ (2018-02-27). "A global network of biomedical relationships derived from text". Bioinformatics. 34 (15): 2614–2624. doi:10.1093/BIOINFORMATICS/BTY114. PMC 6061699. PMID 29490008.
- ^ Robert Leaman; Chih-Hsuan Wei; Alexis Allot; Zhiyong (1 June 2020). "Ten tips for a text-mining-ready article: How to improve automated discoverability and interpretability". PLOS Biology. 18 (6): e3000716. doi:10.1371/JOURNAL.PBIO.3000716. ISSN 1544-9173. PMC 7289435. PMID 32479517. Wikidata Q96032351.
- ^ Wang, Lucy Lu; Lo, Kyle (2020-12-07). "Text mining approaches for dealing with the rapidly expanding literature on COVID-19". Briefings in Bioinformatics. 22 (2): 781–799. doi:10.1093/BIB/BBAA296. PMC 7799291. PMID 33279995.
- ^ Neumann M, King D, Beltagy I, Ammar W (2019). "ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing". Proceedings of the 18th BioNLP Workshop and Shared Task. Florence, Italy: Association for Computational Linguistics: 319–327. arXiv:1902.07669. doi:10.18653/v1/W19-5034. S2CID 67788603.
- ^ Altman RB, Bergman CM, Blake J, Blaschke C, Cohen A, Gannon F, et al. (2008). "Text mining for biology--the way forward: opinions from leading scientists". Genome Biology. 9 Suppl 2 (Suppl 2): S7. doi:10.1186/gb-2008-9-s2-s7. PMC 2559991. PMID 18834498.
- ^ Hanspers, Kristina; Riutta, Anders; Summer-Kutmon, Martina; Pico, Alexander R. (2020-11-09). "Pathway information extracted from 25 years of pathway figures". Genome Biology. 21 (1): 273. doi:10.1186/S13059-020-02181-2. PMC 7649569. PMID 33168034.
- ^ a b Kuhn, Tobias; Royer, Loïc; Fuchs, Norbert E.; Schröder, Michael (2006-01-01). "Improving Text Mining with Controlled Natural Language: A Case Study for Protein Interactions". Lecture Notes in Computer Science. 4075: 66–81. doi:10.1007/11799511_7. ISBN 978-3-540-36593-8.
- ^ Singhal, Ayush; Leaman, Robert; Catlett, Natalie; Lemberger, Thomas; McEntyre, Johanna; Polson, Shawn; Xenarios, Ioannis; Arighi, Cecilia; Lu, Zhiyong (2016). "Pressing needs of biomedical text mining in biocuration and beyond: opportunities and challenges". Database. 2016: baw161. doi:10.1093/database/baw161. ISSN 1758-0463. PMC 5199160. PMID 28025348.
- ^ Hirschman L, Yeh A, Blaschke C, Valencia A (2005). "Overview of BioCreAtIvE: critical assessment of information extraction for biology". BMC Bioinformatics. 6 Suppl 1 (Suppl 1): S1. doi:10.1186/1471-2105-6-s1-s1. PMC 1869002. PMID 15960821. S2CID 5119495.
