Monday, October 13, 2008

WilsonWeb Thesaurus Database

WilsonWeb is a hybrid database with subjects of humanities, social science, education, business, applied science and technology. The thesaurus database is a very useful tool for searchers. For instance, if you are not sure the term indexed in the database, you can search thesaurus database, then you will get the list for all related subjects and related terms with the number of linked records. This is a preliminary search, it will give you some hints to search WilsonWeb with indexing terms.

However, WilsonWeb covers subjects in different disciplines, the same concept could have different meanings in different domains. It would be a huge amount of works to create a hierachic structure in its thesaurus database, or people should call it taxonomy, which will really narrow down the search terms. It's always a challenge for database producers that what kind of thesaurus should be offered to users. Database producers would think cost and effectiveness are the key to solve this problem.

Today, taxonomy has been a main information technology to make search engine more intelligent. Law firms, R & D, and consulting companies have begun building their own taxonomy to enhance the searchability of web search engines, which greatly saves searchers' time with more relevant search results. The problem most people have today is too much information exists, how are they able to find the needed information? Taxonomy could assist companies to organize information and make information easily searchable for users. SLA website and Askus.com are good examples of web content with the benefit of taxonomy.

Monday, October 6, 2008

Metadata Quality Control

Metadata quality control is becoming more important than ever when we deal with batch import. These problems include various versions of author name, unconventional abbreviations, inappropriate data formats, duplicate records, spelling errors, incomplete punctuations, unrecognized characters etc. To control the quality of metadata, we need systematical way to instantly identify and correct those errors; on the other hand, we also need strengthen metadata creation at the beginning.

Recently, DSpace has released an add-on for metadata quality control, which has powerful mass-edit feature, duplicate detection and resolving algorithms. This is a very promising feature for metadata quality control, at least, some systematical methods could be adopted to solve the current dilemma at the time of submission.

Hopefully, those functions, which are available in traditional Integrating Library System, such as, authority control, control vocabulary etc., will be available in Institutional Repository software.

Sunday, September 28, 2008

Metadata Harvesting with MarcEdit

Last time, when I harvested DSpace metadata with MarcEdit, it only lasted around 3 minutes. From the test on yesterday, it went well even though I only did a small collection.

It is important to ensure consistent metadata and access to those digital collections, especially during the period of senior seminars. A simple and clear processing manual should be written to train staff or student assistants how to create metadata, publish theses and harvest metadata with MarcEdit. I am also interested in harvesting metadata with other software. I hope I can hear some good news from Claude in Montreal soon!

Tuesday, September 2, 2008

Back To the Library World

This is a really busy and nice summer. I am glad back to work even though the beginning of the semester is over crowed on campus.

The other exciting is my NITLE Technology Fellowship started in July. Because of the scheduling, I missed the Technology Fellow Workshop on July 21-23 at Southwestern University in Georgetown, TX. But I was able to make it up in August via MIV (Multipoint Interactive Videoconferencing). I met a couple of NITLE Technology Fellows on MIV. Actually, MIV will be my major classroom in the future. I really get excited about what I should teach there ...

Monday, June 30, 2008

RDA Webcasts by Barbara Tillett

Barbara Tillett, the chief officer of the Cataloging Policy and Support Office, Library of Congress,talks about RDA (Resource Description and Access) in two new webcasts. He introduces the background of RDA development, gives us an overview of the new rules, and addresses the next generation cataloging code designed for the digital environment. Will RDA be published in early 2009?

Title 1: Resource Description and Access: Background / Overview
Title 2: Cataloging Principles and RDA: Resource Description and Access

Monday, June 16, 2008

NITLE DSpace User Community Meeting

NITLE DSpace User Community Meeting was held on June 11-12, at University Peget Sound, Tacoma, WA. All the participating institutions have sent their information professionals at the meeting.

It is a great opportunity for participates to share their best practices and also address their concerns. Topics such as institutional repository, digital collection, metadata creation, DSpace Manakin, marketing DSpace etc. have been fully discussed. Most librarians feel that DSpace provides a great opportunity for institutions to create digital collections and scholarly communication on campus wide, but how to effectively promote it on campus and create value-added collection for research will be the key to succeed. NITLE agrees to continue providing such opportunities to facilitate best practices.