{"id":14063,"date":"2016-04-25T16:44:46","date_gmt":"2016-04-25T23:44:46","guid":{"rendered":"http:\/\/nnlm.gov\/pnr\/dragonfly\/?p=14063"},"modified":"2026-02-03T17:11:12","modified_gmt":"2026-02-03T17:11:12","slug":"the-inaugural-pnr-journal-club-data-curation-data-management-and-librarians","status":"archive","type":"post","link":"https:\/\/news.nnlm.gov\/region_5\/the-inaugural-pnr-journal-club-data-curation-data-management-and-librarians\/","title":{"rendered":"The Inaugural PNR Journal Club: Data Curation, Data Management, and Librarians"},"content":{"rendered":"<p><em>This is a guest post by Andrea Harrow, MLS, Good Samaritan Hospital, Los Angeles, CA, about the PNR Journal Club, held over a period of eight weeks from February 8 through March 21, 2016.<\/em><\/p>\n<p>I responded to an invitation sent out on a listserv for involvement in a new journal club&#8211;an opportunity for librarians to discuss issues in data curation and data management. Twelve librarians from across the US also signed up to give it a try.\u00a0 This journal club, hosted by the National Network of Libraries of Medicine, Pacific Northwest Region, met online within Moodle and live via four AdobeConnect sessions. This club followed the\u00a0<a href=\"http:\/\/www.mlanet.org\/p\/cm\/ld\/fid=381\">MLA&#8217;s Discussion Group Program<\/a>\u00a0structure and the new\u00a0PubMed Commons Journal Clubs\u00a0commenting format.\u00a0<!--more--><\/p>\n<p>Before we met live via AdobeConnect, club members introduced themselves on the Moodle site and shared what they hoped to get out of the discussions.\u00a0 We were all interested in the game-changing promise of using big data in health sciences, but, as librarians, were not quite sure how we fit into the big data picture.<\/p>\n<p>The article selection process took place online, using a discussion board in Moodle. We searched out relevant articles, posted these to the forum, then voted on which ones we wanted to discuss. A collection of a few \u201cinspiration\u201d articles steered us toward certain themes. Club members self-nominated to lead the meeting discussion and take notes. \u00a0Using a discussion board setting was a great way to get thematic thought processes flowing.\u00a0 Some of the major themes we discussed on the forum were:<\/p>\n<ul>\n<li>Is evidence-based, guideline-based medicine the enemy of personalized, precision, genomic medicine? \u2013 do they have conflicting priorities, or can they complement each other (<a href=\"http:\/\/blogs.cdc.gov\/genomics\/2014\/02\/13\/is-evidence-based\/\" target=\"_blank\" rel=\"noopener noreferrer\">1<\/a>,<a href=\"http:\/\/informaticsprofessor.blogspot.com\/2015\/05\/is-medicine-precise-enough-to-achieve.html\" target=\"_blank\" rel=\"noopener noreferrer\">2<\/a>)?<\/li>\n<li>Size and significance of data, data biases when data is not adequately translated or translated out of context (<a href=\"http:\/\/informaticsprofessor.blogspot.com\/2016\/01\/biomedical-data-science-needs-measures.html\" target=\"_blank\" rel=\"noopener noreferrer\">3<\/a>).<\/li>\n<li>What is big data? How do big data systems function? Background practical info and overview of the features of clinical big data, algorithms, statistical methods, and software toolkits for data manipulation and analysis, sharing big data, challenges and limitations (<a href=\"http:\/\/www.ncbi.nlm.nih.gov\/pubmed\/25600256\" target=\"_blank\" rel=\"noopener noreferrer\">4<\/a>,<a href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC4214055\/\" target=\"_blank\" rel=\"noopener noreferrer\">5<\/a>,<a href=\"http:\/\/www.icmje.org\/news-and-editorials\/M15-2928-PAP.pdf\" target=\"_blank\" rel=\"noopener noreferrer\">6<\/a>,7).<\/li>\n<\/ul>\n<p>The \u201cUsing Data to Improve Clinical Patient Outcomes\u201d forum, held on March 7, happened to fall on the same day as one of our meetings. Some members attended the forum online or in person. We had selected an article by one of the panelists, Dr. Christopher Longhurst, to discuss in the meeting following the forum.<\/p>\n<p>A summary of our discussions follows.<\/p>\n<p><strong>Week 1<\/strong>:<\/p>\n<p>Cohen B,\u00a0et al.\u00a0Challenges Associated With Using Large Data Sets for Quality Assessment and Research in Clinical Settings. <em>Policy Polit Nurs Pract. <\/em>2015;16(3-4):117-24. PMID:<a href=\"http:\/\/www.ncbi.nlm.nih.gov\/pubmed\/?term=26351216\" target=\"_blank\" rel=\"noopener noreferrer\">26351216<\/a><\/p>\n<p>Facilitator: Laura Zeigen. Recorder: Erin Foster.<\/p>\n<p style=\"padding-left: 30px\">The\u00a0article\u00a0addresses\u00a0the challenges of\u00a0interfacing biomedical big data\u00a0to answer\u00a0clinical and epidemiological\u00a0research questions, and how these challenges\u00a0were overcome in\u00a0the construction\u00a0of a &#8220;data-mart&#8221;\u00a0of\u00a0medical, financial (i.e., billing), and demographic patient information in\u00a0an\u00a0academically affiliated health-care network. This article presents seven challenges identified by the National Institutes of Health (NIH) Big Data to Knowledge (BD2K) initiative and recommendations\u00a0for\u00a0overcoming those challenges\u00a0 based on the project experience. \u00a0\u00a0One of the BD2K identified\u00a0obstacles of biomedical big data\u00a0is the organization, management, and processing of data. The varied ways in which clinical data is collected makes it difficult to validate for consistency as well as discover for re-use.<\/p>\n<p style=\"padding-left: 30px\">Another obstacle identified in the paper was the challenge\u00a0of training researchers to use data effectively. The\u00a0article primarily\u00a0discusses this in the context of hiring a data manager and the difficulty of identifying a person that has the necessary combination of skills or the job. For members, this raised questions about what kinds of skills librarians should consider developing in order to be valued, useful members of these teams.<\/p>\n<p style=\"padding-left: 30px\">Discussion revolved around several themes:\u00a0challenges librarians face in identifying\/occupying data science roles, the potential of promoting existing\u00a0librarian\u00a0skills and cultivating new skill sets, and the importance of networking and professional development.\u00a0Journal club members cited librarian training in organizing and tagging\/indexing\u00a0data as well as enhancing data by connecting it to the published literature.\u00a0 Several members called attention to the\u00a0lack of published librarian contributions to\u00a0data science, especially since this article was authored by a multidisciplinary team.<\/p>\n<p><strong>Week 2:<\/strong><\/p>\n<p>Marshall DA, et al. Transforming Healthcare Delivery: Integrating Dynamic Simulation Modelling and Big Data in Health Economics and Outcomes Research. Pharmacoeconomics. 2016 Feb; 34(2):115-26. PMID: <a href=\"http:\/\/www.ncbi.nlm.nih.gov\/pubmed\/?term=26497003\" target=\"_blank\" rel=\"noopener noreferrer\">26497003<\/a>.<\/p>\n<p>Facilitator<strong>: <\/strong>\u00a0Ann Gleason.\u00a0 Recorder: Suzanne Fricke.<\/p>\n<p style=\"padding-left: 30px\">This article presents the idea of Dynamic Simulation Modeling (DSM), a computerized mathematical modeling used for several years in other disciplines, to develop meaningful insights from big data in healthcare environments. Applications of DSM to health care are complex due to de-identification requirements and harmonization of health care data sources (EMR, insurance, wearables, social media and large inter-organizational datasets). We are optimistic that there will be an ongoing need for librarian skills in organizing data, taxonomy development, naming conventions, metadata creation, and retrospective adaptations to changing terminology.<\/p>\n<p style=\"padding-left: 30px\">We noted the increasing need for skills in data science, machine learning and programming. Data cleaning and de-identification remains highly labor intensive in the health care setting. Although statistical software (SAS, R, MatLab, etc.) may be provided to users by libraries, we had difficulty envisioning what these mathematical models might look like and what novel software they might require to run.\u00a0 We discussed the need for tools such as data animation models to show students the steps and skill sets required for healthcare big data.<\/p>\n<p style=\"padding-left: 30px\">Big data increases the need for consumer resources centered on informed consent, particularly in relation to <em>universal consent<\/em>. The Henrietta Lacks case and PatientsLikeMe were mentioned as differing consumer perspectives on the unanticipated long-term use of patient data on the one hand, and the symbiotic relationship between patients and researchers looking for free access to medical advice, treatment and data on the other.<\/p>\n<p><strong>Week 3: <\/strong><\/p>\n<p>Panahiazar M, et al. Empowering Personalized Medicine with Big Data and Semantic Web Technology: Promises, Challenges, and Use Cases. Proc IEEE Int Conf Big Data. 2014 Oct; 2014:790-795. PMID: <a href=\"http:\/\/www.ncbi.nlm.nih.gov\/pubmed\/?term=25705726\" target=\"_blank\" rel=\"noopener noreferrer\">25705726<\/a>.<\/p>\n<p>Facilitators: Erin Foster and Carol Perryman.\u00a0 Recorder: Lynly Beard.<\/p>\n<p style=\"padding-left: 30px\">The themes of this article were raised in past club discussions: 1)\u00a0examining clinical co-morbidities and genomics to arrive at personalized patient treatment and 2) handling the volume of data with new technologies, to make it &#8220;smart data&#8221; and use it for analysis purposes.\u00a0 Creating data sets allows for uniform analysis, contextualization, and moves data along to the goal of personalized medicine.<\/p>\n<p style=\"padding-left: 30px\">Three interconnected use cases revolved around &#8220;heart failure&#8221;. \u00a0With these, the authors introduced new technologies used to transform big data into smart data. Case #1 used <strong>Hadoop<\/strong>, a tool that can handle large volumes of data by breaking them down into smaller batches processed at several nodes, as well as the Pig programming tool. Case #2 used the <strong>UMLS\u00a0Metamap<\/strong> tool, which when used in conjunction with Hadoop and Pig, reduced processing time for 10 million search queries from 40 days to 2 days. Case #3 then used <strong>Kino<\/strong> to add\u00a0metadata.<\/p>\n<p style=\"padding-left: 30px\">These tools were new to many in the group, and there was a desire for further information.\u00a0 Because MetaMap is an NLM tool, participants were curious about training opportunities. \u00a0We again discussed using a critical evaluation framework for all articles. \u00a0This preprint article would have benefitted from expanding the use cases and putting them in context.<\/p>\n<p><strong>Week 4: <\/strong><\/p>\n<p>Longhurst CA, et al. A &#8216;green button&#8217; for using aggregate patient data at the point of care. Health Aff (Millwood). 2014;33(7):1229-35. PMID: <a href=\"http:\/\/www.ncbi.nlm.nih.gov\/pubmed\/?term=25006150\" target=\"_blank\" rel=\"noopener noreferrer\">25006150<\/a>.<\/p>\n<p>Facilitator: Andrea Ball. Recorder: Laura Zeigen.<\/p>\n<p style=\"padding-left: 30px\">Evidence-based medicine (EBM) has traditionally been based on randomized control (RCT) data. However, RCTs are expensive, time consuming, and not easily generalizable.\u00a0 Longhurst <em>et al.<\/em> introduce the concept of a \u201ccontinuous learning health care system\u201d (coined from Institute of Medicine) through utilization of the electronic health record (EHR).\u00a0 The \u201cgreen button,\u201d placed within a patient EHR would help clinicians find similar patients and provide support for patient care decisions in the absence of evidence. Precursors to the \u201cgreen button\u201d idea are the \u201cblue button\u201d from the VA system (for beneficiaries to obtain health care information in a consolidated way) and \u201cinfobuttons.\u201d (Cimino, 2013).<\/p>\n<p style=\"padding-left: 30px\">Challenges with implementing the green button include the necessary policies and incentives to be in place, HIPAA\/privacy, IRBs, informed consent, and visualizing the results in a meaningful way. \u00a0The authors suggest health care systems initially approach integration of the button as a qualitative improvement process.<\/p>\n<p style=\"padding-left: 30px\">Patient preferences are the \u201cthird leg of the stool\u201d for EBM.\u00a0 If patient preferences or values were included in the patient record, similar patients with similar care preferences could be identified. \u00a0This is the first article we have read that discusses the obligation placed on the patient to contribute patient-generated data.\u00a0 Might there be levels of participation to which a patient could agree?\u00a0 Issues of informed consent, privacy and related ethical concerns could be part of what librarians help bring to the table. Another potential role for librarians, here, is mapping MeSH terms to SNOMED and other vocabularies.<\/p>\n<p>References:<\/p>\n<ol>\n<li><a href=\"http:\/\/blogs.cdc.gov\/genomics\/2014\/02\/13\/is-evidence-based\/\">http:\/\/blogs.cdc.gov\/genomics\/2014\/02\/13\/is-evidence-based\/<\/a><\/li>\n<li><a href=\"http:\/\/informaticsprofessor.blogspot.com\/2015\/05\/is-medicine-precise-enough-to-achieve.html\">http:\/\/informaticsprofessor.blogspot.com\/2015\/05\/is-medicine-precise-enough-to-achieve.html<\/a><\/li>\n<li><a href=\"http:\/\/informaticsprofessor.blogspot.com\/2016\/01\/biomedical-data-science-needs-measures.html\">http:\/\/informaticsprofessor.blogspot.com\/2016\/01\/biomedical-data-science-needs-measures.html<\/a><\/li>\n<li><a href=\"http:\/\/www.ncbi.nlm.nih.gov\/pubmed\/25600256\">http:\/\/www.ncbi.nlm.nih.gov\/pubmed\/25600256<\/a><\/li>\n<li><a href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC4214055\/\">http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC4214055\/<\/a><\/li>\n<li><a href=\"http:\/\/www.icmje.org\/news-and-editorials\/M15-2928-PAP.pdf\">http:\/\/www.icmje.org\/news-and-editorials\/M15-2928-PAP.pdf<\/a><\/li>\n<li>http:\/\/www.niso.org\/apps\/group_public\/download.php\/15375\/PrimerRDM-2015-0727.pdf<\/li>\n<li>Cimino JJ, Li J<em>. <\/em>Sharing infobuttons to resolve clinicians\u2019 information needs. AMIA Annu Symp Proc. 2003;2003:815<\/li>\n<\/ol>\n<p>Special thanks to the members of this journal club (Andrea Ball, Lynly Beard, Erin Foster, Suzanne Fricke, Ann Gleason, Mary Anne Hansen, Andrea Harrow, Ayaba Logan, Carol Perryman, and Laura Zeigen). Emily Glenn, Community Health Outreach Coordinator, NN\/LM PNR, was the club moderator.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>This is a guest post by Andrea Harrow, MLS, Good Samaritan Hospital, Los Angeles, CA, about the PNR Journal Club, held over a period of eight weeks from February 8 through March 21, 2016. I responded to an invitation sent out on a listserv for involvement in a new journal club&#8211;an opportunity for librarians to&#8230; <a href=\"https:\/\/news.nnlm.gov\/region_5\/the-inaugural-pnr-journal-club-data-curation-data-management-and-librarians\/\">Read More &raquo;<\/a><\/p>\n","protected":false},"author":22,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_monsterinsights_skip_tracking":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[8,5],"tags":[],"class_list":["post-14063","post","type-post","status-archive","format-standard","hentry","category-technology","category-training-education"],"jetpack_sharing_enabled":true,"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/posts\/14063","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/users\/22"}],"replies":[{"embeddable":true,"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/comments?post=14063"}],"version-history":[{"count":3,"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/posts\/14063\/revisions"}],"predecessor-version":[{"id":17667,"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/posts\/14063\/revisions\/17667"}],"wp:attachment":[{"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/media?parent=14063"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/categories?post=14063"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/news.nnlm.gov\/region_5\/wp-json\/wp\/v2\/tags?post=14063"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}