Advanced Search
Journal search
Journal cover: Library Hi Tech

Library Hi Tech

ISSN: 0737-8831

Online from: 1983

Subject Area: Library and Information Studies

Content: Latest Issue | icon: RSS Latest Issue RSS | Previous Issues

Options: To add Favourites and Table of Contents Alerts please take a Emerald profile

Previous article.Icon: Print.Table of Contents.Next article.Icon: .

Heuristics for identification of bibliographic elements from title pages

Document Information:
Title:Heuristics for identification of bibliographic elements from title pages
Author(s):Durga Sankar Rath, (Lecturer in the Department of Library and Information Science, Ravindra Bharati University, Kolkata, India), A.R.D. Prasad, (Associate Professor, Documentation Research and Training Centre, Indian Statistical Institute, Bangalore, Karnataka, India)
Citation:Durga Sankar Rath, A.R.D. Prasad, (2004) "Heuristics for identification of bibliographic elements from title pages", Library Hi Tech, Vol. 22 Iss: 4, pp.389 - 396
Keywords:Bibliographic systems, Cataloguing, Classification schemes, Data handling, Information operations
Article type:Research paper
DOI:10.1108/07378830410570494 (Permanent URL)
Publisher:Emerald Group Publishing Limited
Abstract:This paper presents a methodology for automatic identification of bibliographic data elements from the title pages of books. Also enumerates the various steps like scanning the title pages, running optical character recognition (OCR) software, generating HTML files out of title pages and applying heuristics to identify the bibliographic data elements. Much of the paper deals with the surveys undertaken to analyze the characteristics of various bibliographic descriptive elements like title, author, publisher and other elements. The first survey deals with the sequence of the bibliographic data in the title pages. The second survey deals with the font size, font type and the proximity of each bibliographic element on the title pages. The survey results are then used to develop heuristics, in order to develop a rule-based expert system which can identify the bibliographic elements on the title pages. The results of the system are presented, along with problems encountered.

Fulltext Options:



Existing customers: login
to access this document


- Forgot password?
- Athens/Institutional login



Downloadable; Printable; Owned
HTML, PDF (277kb)Purchase

To purchase this item please login or register.


- Forgot password?

Recommend to your librarian

Complete and print this form to request this document from your librarian

Marked list

Bookmark & share

Reprints & permissions