Showing posts with label Google. Show all posts
Showing posts with label Google. Show all posts

Thursday, May 22, 2008

Schottlaender's "On the Record" presentation

Additional reports and presentations from the Spring Assembly May 7, 2008.

Brian Schottlaender (UCSD) "On the Record"; The Library of Congress Working Group on the Future of Bibliographic Control

The Library of Congress, in response to the evolving information and technology environment, convened the Future of Bibliographic Control Working Group to examine the future of bibliographic description in the 21st century. As a member of the working group, Schottlaender will discuss the group’s final report and the implications and ramifications of the report or the UC libraries.

Referred to in presentation:
On the Record: Report of the Library of Congress Working Group on the Future of Bibliographic Control
presented: January 9, 2008

Thomas Mann. “'On the Record’ but Off the Track” - a response on behalf of the Library of
Congress Professional Guild

LC’s Cataloging Policy and Support Office has issued decisions regarding LCSH
http://www.loc.gov/catdir/cpso/pre_vs_post.pdf

Thursday, May 8, 2008

Presentation slides from Stephen Abram talk "Heading for the 3.0 World"

Stephen Abram (SirsiDynix): Heading for the 3.0 World: Technologies and Behaviors to Watch (PPT 26.4Mb) (PDF 8.3 Mb)

Abstract: Can academic libraries be more open? Can we be more open to our scholars, our researchers, our learning communities, to new technologies? Can we be more open to change? How? Are there technologies that we should be trying and piloting to see if they improve the library's mandate? Which ones are worth investigating? What are the emerging learning technologies? Are there different and improved ways to enhance our organization's missions? Can we enhance our research and learning communities and attract more funding and use? What about books, OPACs, databases and interfaces? What changes are happening here? Stephen Abram is an inveterate library watcher and strategic technology futurist for libraries. In this session, he shares the top technologies that we should think about 'playing' with while finding a way to make our libraries more open to our learning, publishing and research communities. Can we drive quicker adaptation to change in our own library culture? He will end with five suggestions about how to have fun with change and technology adoption.

Slides will be linked from Stephen's Lighthouse blog (http://stephenslighthouse.sirsidynix.com/)

Friday, May 2, 2008

Background on Draft Report of the Working Group on the Future of Bibliographic Control

Draft Report of the Working Group on the Future of Bibliographic Control

http://www.loc.gov/bibliographic-future/news/draft-report.html

"In reading the report, you will note that its findings and recommendations are structured around five central themes:

1. Increase the efficiency of bibliographic production for all libraries through increased cooperation and increased sharing of bibliographic records, and by maximizing the use of data produced throughout the entire "supply chain" for information resources.

2. Transfer effort into higher-value activity. In particular, expand the possibilities for knowledge creation by "exposing" rare and unique materials held by libraries that are currently hidden from view and, thus, underused.

3. Position our technology for the future by recognizing that the World Wide Web is both our technology platform and the appropriate platform for the delivery of our standards. Recognize that people are not the only users of the data we produce in the name of bibliographic control, but so too are machine applications that interact with those data over the network in a variety of ways.

4. Position our community for the future by facilitating the incorporation of evaluative and other user-supplied information into our resource descriptions. Work to realize the potential of the FRBR framework for revealing and capitalizing on the various relationships that exist among information resources.

5. Strengthen the library profession through education and the development of metrics that will inform decision-making now and in the future. "

"The period for public comment on the report is open until December 15, 2007. Comments can be submitted via the Web site at http://www.loc.gov/bibliographic-future/contact/. Electronic submission of comments is encouraged. "


From LJ Academic Newswire
http://www.loc.gov/today/pr/2007/07-219.html

LC: Draft Report on Bibliographic Control To Be Released Nov. 13, 2007

For a year, the library world has been watching to see what the Working Group on the Future of Bibliographic Control, convened by the Library of Congress (LC), will say about the future of bibliographic description given the increasing reliance on web-based searching and electronic information resources. The wait is nearly over. LC officials said today that a draft report will be presented to LC managers and staff at 1:30 p.m. EST on Nov. 13, along with a live webcast. A comment period will follow and last until Dec. 15.

Even before the announcement, however, American Library Association (ALA) President-elect Jim Rettig, in testimony Oct. 24 before Congress, expressed concern that LC not move too precipitously. Rettig, university librarian of the Boatwright Memorial Library, University of Richmond, VA, told the Committee on House Administration, that ALA "strongly recommends that the Library of Congress return to its former practice of broad and meaningful consultation prior to making significant changes to cataloging policy." Rettig said he hoped LC fully "understands the impact" that its decisions have on other libraries, noting that LC bibliographic records "are accepted without editing by thousands of libraries of all types and sizes throughout the world to facilitate an individual's access to library resources."

He added, "Inevitably, on the Internet, with its huge and ever-increasing amount of digital information, general search engines must be relied upon. And, in years to come, there may be far more sophisticated search engines. But we are certainly not there now. The consumers of the Library's cataloging products must continue to rely on the traditional cataloging services in order to meet the needs of their users…. Further, unilateral and sudden changes to cataloging practice initiated by the Library of Congress and others severely and negatively affect citizens' ability to find answers in libraries and elsewhere."

Information on the Working Group and its findings is available at www.loc.gov/bibliographic-future/

British Library response to the Library of Congress Working Group on. the Future of Bibliographic Control http://www.bl.uk/services/bibliographic/pdf_files/bl_response_lcwgfbc(final).pdf

About the speaker: Brian E. Schottlaender is the Audrey Geisel University Librarian at the University of California, San Diego. Prior to joining UC San Diego in 1999, his career in libraries included positions at the California Digital Library, UCLA, the University of Arizona, Indiana University, and int he European book trade. In 2008, Schottlaender was appointed Secretary of the Board of Directors of The Center for Research Libraries (CRL), a consortium of North American universities, colleges, and independent research libraries that acquires and preserves traditional and digital resources for research and teaching. In addition, he has been elected to the members Council of OCLC, a nonprofit, membership, computer library service and research organization that serves more than 60,000 libraries in 112 countries internationally, and serves on the Steering Committee for the Coalition of Networked Information (CNI). He was president of the Association of Research Libraries (ARL) in 2006.

Thursday, November 15, 2007

questions for the presenters -- Robin Chandler and John Kunze

Question: OCR -- what's the success of OCR, error-wise? What kind of editing do you have to do?

JK: we looked at the degradation of OCR over time vs compression -- but doesn't have any data on the average error rate per book. (An idea: have library school students go through and correct pages as part of learning about OCR).*

Question: Languages -- apparently there are certain languages that don't OCR well?

A: German and CJK (Chinese, Japanese, Korean scripts) and Greek are problematic -- but Google etc. don't have the tools to index these scripts either. There's a product called Abireader (?) that the IA is using for Russian.

Question: Google had controversy that they were western-based, U.S. centric -- but BHG was trying to find an Indian publication (in English) that wasn't available anywhere the other day. Is there a push to move digitization beyond just the libraries we have heard about and move it into collections we couldn't get any other way?

Answer: The answer is yes -- Google especially has moved into Europe and Japan, but not India yet. RC thinks they are pushing for to get out more.

Q: is there room for bibliographers to give guidance to Google? Can we say, "do these, they're rare and public domain?"

A (RC): for our UC piece, the answer is yes

Michigan is the only library that has put content from the digitization up, and they have built a large rights database of who can use what

Q: are there any restrictions on our use of the project?

A (RC): Google has three different versions of what you can view -- title, snippet, full view

MS & OCA is only scanning material in public domain

UC restrictions... we can't make copyrighted material available either -- PD material we can share (obviously) -- restricted in the percentage of public domain material we can share via google (???)

content contract -- we can't allow the content to be "indexed or downloaded by a commercial service"

RC: Google is trying to follow all the laws in all the countries.. .

Q: who's doing the work on all these orphaned works to find out if they really are orphaned?

A: OCA, MS and Google are all interested in it, but OCA is doing a lot of the work

The Boston Library consortium has said that they will digitize things if someone requests them

Q: what is the speed of searching all these digitized books -- esp if they put full text into worldcat.

A: OCLC and Google have not finalized how they will put links to books into worldcat.

Q: will this material be accessible to anyone, or will it only be accessible through a proxy server (e.g. if I'm helping a non-UC student)


A (RC):
in the pilot, when there's a link from the record to the item at Google or MS, anyone who can get into the catalog can see that book -- restricted by copyright

what about copyrighted works that we own -- not addressed

Q: limitations that we've agreed to in our contract -- are there any restrictions

A (RC): yes. what do we want to do?
we have not agreed to restrictions beyond the copyright issues etc.

Q: Can they copyright the digital form that they have made of our books?

A: (RC) -- I don't think so...

Q: who is checking the quality of the digitized works? What are they doing?

A: (RC) - Google puts a lot of effort into checking quality and quality algorithms. JK: At CDL we do a bunch of format checking to make sure the files are well formed. RC: we don't have the staff at CDL to go through every page, beyond the files.. but the vendors are interesting in getting error reports from individuals, though there's some question of how to do that and how do rescans/insert pages, etc. Might fall on the library to help correct errors...

BHG: Google is no longer double-scanning, they're comfortable with the number of errors they are getting now.

Q) how are books scanned?
RC: in both processes, the scanning is manual -- it's people turning pages. But the processes are different... the artificial intelligence part comes into error correction.

Google has several scanning centers around the country, but they are not outsourced. They are not using automatic machines..

q) what about fold-out maps, and other rare materials?

RC) None of the projects can do folios/large format. MS & IA -- when they scan, they have been skipping books with fold-out maps (but tracking on those lists). We've been working with MS/IA, telling them we'd like to get those foldouts scanned -- so IA has been working on trying to get them done in future in an elaborate process.

Google is scanning the book, but not scanning the foldout.

Q: Is there a problem with mislaying titles?

A: (RC) -- They have not lost any books so far.. with Google, there's one that has gone missing recently, and apparently that's the only book they have lost in all their 27 sites. They have "shipment reconciliation statements" from NRLF to Google -- they'are on it all the time.

Q: do you have any interest in or pressure from faculty on what gets done?

A: (RC) -- Honestly, we haven't done enough to really engage the faculty yet. Re the access piece -- we haven't really done much; Google has been talking to digital humanities centers around the country, which is great, but that's also our job as well.

Q: can you give us a preview of your digitization feedback group and what you plan to be doing, and how you plan to get feedback?

A: (RC) -- the challenge will be that any proposal that comes forward needs to have been signed off by a UL on campus etc etc.

* As a fairly recent grad, I have to say this doesn't sound like much fun...

Questions about OCR

Quality of OCR and what kind of editing is required? They looked at the degradation of quality and OCR performed better with fewer errors and then got worse with over compression. It depends on the original book or item. There are no efforts right now for corrections right now. Perhaps we could assign library school students to correct pages 25-30 for homework and over the years we'll get a lot more items corrected.

Google folks were talking about things that might be better than OCR. A lot of foreign language items are not getting useable with OCR, what can we do? OCR does vary greatly by language but these days results are gettting better. We have to prioritize languages with the usage by our patrons. Germans, Islamic languages & Greek do have a lot of problems. CJK is making a lot of progress. Abbey 8 is Google's OCR and it's in theory moving along pretty quickly. Google is working on that. They are commercial entities and they are looking at where they can make money and so Google is very interested in Asia right now.

Google is expanding into Europe and Japan in bringing their collections into the fold but India has not yet been approached. They tried with France, but they are doing their own thing.

"jewels of the collection"

Chandler said that some of the jewels of the digitized collections include special collections, Bancroft Library, classic mathematics, children's books, cookbooks...

she showed some slides of some of these. Of course, there aren't links in our catalogs yet for these -- she suggested that when these come the project will become a little more "real" for all of us.

Chandler referenced this blog: http://landscape.blogspot.com/ which talks about the effect of google books on her work as a scholar.