July 4, 2009
| Bio-IT World > Full Text Firewall
Full Text Firewall


Oct. 16, 2006 | One of the biggest stumbling blocks to the success of text mining remains the firewall that surrounds full-text archives. "It is the restricted access to the full text of papers and to citation information, rather than the technology, that is currently the greatest limitation, despite some encouraging open-access initiatives," EMBL bioinformaticist Peer Bork and colleagues wrote in Nature Reviews Genetics earlier this year.

Access to the full text often requires exorbitant subscriptions to publishing powerhouses such as Elsevier, Wiley, and others. The open-access movement, which fueled the launch of the Public Library of Science in 2002, has persuaded some publishers to partially open their archives - but most new papers remain off limits to non-subscribers.

Some of QUOSA's customers have done benchmarking, including one whom MacKenzie says compared full-text versus abstract text mining for protein-protein and protein-drug reactions, and retrieved 25-100 times more useful factoids out of full-text sources. Would more publishers going the open-access route reduce the need for QUOSA? "Open access is just another silo through which you get journal access," says MacKenzie. "The better access there is to full articles, the more people can take advantage of what we do."

Nature Publishing Group (NPG) recently announced an interesting experiment - an effort to enhance machine access to full text literature by proposing a standard content annotation called the Open Text Mining Interface (OTMI), which was first presented by NPG web publishing director Timo Hannay at Bio-IT World's 2006 annual conference.

The XML format of OTMI reorders each paper's sentences alphabetically, rendering the product unreadable to humans while allowing full-text searching of intact sentences. It's what Hannay calls "a potential compromise between business needs and open access." Hannay says he hopes all publishers will adopt OTMI or a similar standard to open up the entire literature for text mining.

One fan is Tim O'Reilly, who says, "It immediately struck me as 'slap your forehead brilliantly obvious...' I love the cleverness of this approach, which lets machines make use of the content in ways that human readers can't. I like it. You might consider it a "copyright hack.' "

Hannay says there's been considerable early interest from other publishers and text-mining researchers. This fall, NPG will roll out OTMI files for much of the Nature archives and encourage people to play with it. NPG is also launching a collaborative website to open up development of the specification and share tools. "Ideally, we want OTMI to become a de facto community standard (and perhaps in time a more formal standard). NPG has no desire to 'own' it," says Hannay. -- K.D.

Return to main article.

 

 

Click here to login and leave a comment.  

0 Comments

Add Comment

Text Only 2000 character limit

Page 1 of 1

White Papers & Special Reports

thomson reuters image
Biomarkers: An Indispensible Addition to the Drug Development Toolkit
Examining the Potential of Biomarkers
Sponsored by Thomson Reuters

Biomarkers are becoming an essential part of clinical development. In this white paper, Thomson Reuters provides insight from experts in industry and academia, and explores the role of biomarkers as evaluative tools in improving clinical research and the challenges this presents.

Discover the potential of biomarkers to:

  • Improve decision making
  • Accelerate drug development
  • Reduce development costs


BlueArc_Scientific Data
Scientific Data Lifecycle Management: Preparing for Storage in an Uncertain Future
Sponsored by BlueArc

Managing vast and overwhelming streams of gene sequencing data today requires ultra-high performance systems and processes. With continued rapid advancement and improvements in gene sequencing, expect tomorrow’s instruments to output quantities of genomic information that will dwarf current levels. Help your organization maintain data control and prepare for the future of sequencing through this informative paper that discusses:

  • The information technology challenges of gene sequencing
  • “Intelligent” methods for data management and customization
  • System survival tips... Deciding what data to keep or delete
  • New tools to keep scientists ahead of impending data torrents


SAS Managed image
Managed Innovation, Assured Compliance
Developing, executing and managing the transformation, analysis and submission of clinical research data with SAS® Drug Development
Sponsored by SAS
Get better products to market faster. Download this white paper to discover the top ten challenges facing life science executives and how to overcome them. See how SAS Drug Development transforms clinical data into true innovation.


Life Science Webcasts & Podcasts

Presented by Trade Commission of Spain

Spain Biotech: An Engine for Economic Change 

TCS podcastDiscover how Spain is focusing on biotechnology to be an engine for economic change through gradual internationalization, development and technology transfer.

Regional governments are actively investing in public and private biology research and promoting the creation of knowledge-based companies. Spain’s human capital combined with aggressive investment in biotech research and infrastructure has led to the creation of bio-clusters.

Today, there are nearly 700 Spanish companies engaged in biotechnology, with almost 50 percent growth in funding devoted to research. In fact, spending on internal R & D in biotechnology has grown 46 percent and is close to 300 million Euros.

Access the podcast 

 



More Podcasts

Job Openings

saic_logo

MANAGER, SCIENTIFIC COMPUTING & PROGRAMMING
(Bioinformatics Manager)
SAIC-Frederick, Inc has an exciting opportunity for a Manager, Scientific Computing & Programming - Core Genoytyping Facility in Gaithersburg, Maryland.  In this role, you will lead the Bioinformatics & Analysis Group.
Master’s or equivalent required.  PhD preferred. Six years experience in development of scientific programs in high-performance computing environment including five years supporting scientific research in computational chemistry, biology, or genetics, & two years supervisory experience.  View complete job posting & apply: www.saic-frederick.com. Position #146945.

For reprints and/or copyright permission, please contact The YGS Group, 1808 Colonial Village Lane, Lancaster, PA;

(717) 399-1900 ext. 125, or via email to Ashley.Zander@theYGSgroup.com.