Showing posts with label metadata. Show all posts
Showing posts with label metadata. Show all posts

Monday, November 30, 2009

Who Should Create Metadata for Online Submissions in DSpace II?

What happens at small academic/research libraries? At some libraries, metadata librarians or archivists create metadata for a variety of collections if metadata standards have been established. However, this is not only a huge amount of work, but also needs professionals to do the job. Today, to facilitate metadata creation, metadata librarians are seeking for batch loading or auto-generated metadata to provide access to digital contents with the benefit of technologies. This is an emerging challenge for Metadata or Digital Initiative Librarians. When libraries migrate a collection to a new platform, batch loading metadata for the collection is more efficient and effective, especially if the collection is not created from the scratch. This also needs metadata librarians to map metadata in one system to those in another with their expertise.

Some people think the library can use students or paraprofessionals to do the job. I would say Yes and No. As you might notice, the purpose of metadata creation is to let users easily find the information they need. If a person doesn't understand the philosophy of the information retrieval, how can he/she know to create the right access for users? However, if metadata librarians can set up some procedures, teach them some of the processes, then they would be a great help to metadata librarians. For instance, if metadata librarians set up batch loading form, students or paraprofessionals will prepare the basic metadata form first, then metadata librarians can work on the form and batch load metadata into DSpace.

Therefore, there is no rules for this. Librarians should allocate tasks with the collaboration of the personnel in the libraries. The bottom line is to facilitate metadata creation for digital collections.

Tuesday, November 17, 2009

Who Should Create Metadata for Online Submissions in DSpace I?

Since the Miller Library started to deposit seniors theses in the college digital repository DSpace in 2007, librarians have been wondering who is the appropriate person to create metadata for theses. Ideally, metadata librarians can do this with best knowledge they have. However, most submitted theses need original cataloging. Usually, there is only one metadata /cataloging librarian in a small academic or research library to do this type of original work. The amount of electronic theses received by the library each year is far beyond what a metadata librarian can handle in a timely manner. Especially, the metadata librarian has to design metadata models for different types of electronic collections, and facilitate access for users to easily search information in Dspace. So what are the possible solutions?

At large university libraries, students submit theses and dissertations online. In this submission process, a senior needs to create descriptive metadata for his own work, such as title, author, keywords, abstract, table of contents etc. After the student submits his/her thesis, the metadata librarian will review the submission. If the thesis passes the review, the librarian will publish it right away. If the metadata librarian finds out inappropriate metadata in the submission, s/he will not publish it until errors are fixed.

Monday, August 31, 2009

Building an Institutional Repository for Your Instutution

Institutional repositories (IRs) have been successfully populated at higher education institutions, where users not only get open access to those scholarly publications, but also create collections. It opens the door to share research information within or beyond the community. IRs are also extremely useful at research companies.

Research companies could set up different communities to share information at different levels. This will reduce duplicate paperwork, such as lab reports, lab records and datasheet. It also minimize the requests for the same information, and leads to a green business environment. The easily customized workflow could be designed to facilate researchers to deposit thier data in a moment. IRs could also serve as a platform for record management. The descriptive metadata and administrative metadata can be shared or transferred as a part of management records.

Currently, some open source software, such as DSpace, Greenstone, Fedora, are widely used at academic libraries. The other commercial products, such as CONTENTdm and Inmagic Presto, are also used in business environment. How to create an institutional repository and promote it in your community? The article, Building an Instutional Repository at a Liberal Arts College, might give you some thoughts and inspiration.

Sunday, April 5, 2009

Metadata Workshop on NIS Camp in June 2009

Self-created digital collection has become an effective means for libraries to create knowledge, preserve and share archival and research information. In order to provide easy online access to these research materials for users, librarians and information professionals use metadata to organize and describe information, and make these resources online searchable. How to create effective metadata for digital collections? Jin Xiu Guo will offer a workshop on NITLE Information Services Camp at Smith College on June 4, 2009. This workshop is for everyone who wants to know about metadata and likes to explore knowlege management in digital age. People who are interested in the workshop could visit NIS Camp.

Wednesday, January 28, 2009

"Author vs. User Tagging" on Journal of Library Metadata

With the increasing application of social tagging technology in web 2.0, librarians have applied social tagging in online library catalog as an additional search entry for users. Now scientists, attorneys and technology consultants start to tag web contents as subject experts to provide such convenience. This user tagging is becoming an acceptable access tool for researchers. But what are the differences between author-supplied metadata (endo-tagging) and user- supplied metadata (exo-tagging)? An interesting article by Heather Lea Moulaison on Journal of Library Metadata, (2008, vol. 8, issue 2 p101-111, 11p) has a critical review on this issue.

Journal of Library Metadata focuses on emerging issues about all aspects of metadata applications in today's digital libraries. Haworth's Journal of Library Metadata, now published by the Taylor & Francis Group, is seeking a new editor. Any interested professionals with sufficient credentials who might like to take on this task can contact Bill Cohen, the publisher at bcohen7719@aol.com.

Tuesday, January 6, 2009

International Standard Collection Identifier (ISCI)

International Standard Collection Identifier is under review now. The purpose of the document is to make various collections and fonds to be identified by a system in a systematical way. An identifier is an important metadata element, so far, there is no standard way to construct an identifier. Different entities have their own ways to create identifiers.

With the appearance of metasearch engine, an identifier has played an important role in locating information. Especially electronic resources have been exponentially increased in a recent decade, an identifier is usually generated by local practices. There is no systematical way to guide local practices to formulate an consistent identifier. Now more and more digital libraries have created digital collections including digital archives, descriptive metadata are used to describe those resources. An identifier is a key element of descriptive metadata. Today, knowledge sharing is an effective learning process. Knowledge sharing can not be alive on its own, metadata sharing and exchange inevitably accompany with knowledge sharing.

When a variety of organizations use different local identifiers, the metasearch engine has to do duplicate searching because of irregular formation of the identifiers. This irregularity greatly reduces search engine efficiency. To increase search effectiveness, we need to build standard identifiers for collections and fonds, to facilitate global metadata exchange. The proposed standard way is:

ISIL:Collection identifier string

ISIL is the identifier for the organization, the collection identifier may contain up to 16 characters.

e.g. FI-Ht:Up Helsinki University Theology Library, Psychology of religion collection. (example from ISO/CD 27730)

If ISCI becomes a standard, it will greatly reduce duplicate detection of a search engine, users will be able to identify collections and fonds through ISCI.

Monday, December 8, 2008

Dublin Core One-to-One Principle

In Dublin Core metadata schema, the one-to-one principle refers to one metadata description is only for one resource. For instance, description for a digital image of Mona Lisa can not be regarded as same as the original painting. However, in most practices, it's difficult to just make a straight line of it.

When we create metadata to describe a resource, such as a digital image, or an analog object, we need to consider users' requirements. From users perspective, we want to give the information they are looking for; metadata creators should have the capability to identify the key information need. For example, when a metadata creator describes the date of an image of Mona Lisa digitized from an original painting, s/he should think about what users really want to know here. In most case, users are interested in the original date of the painting instead of the image. If metadata creators give the digitization date of the image, it would be less satisfied users' interest.

However, in the above example, if the original created date of the painting is provided in the metdata description instead of the digitization date, it would conflict with the one-to-one principle. Therefore, we need to use our best judgment to create metadata meaningful for users, rather than just follow straight rules and miss the information users need.

Monday, November 24, 2008

RDA Constituency Review

RDA is up for review again. People who are interested in RDA could submit your comments by February 2, 2009. RDA (Resource Description and Access) will be the general guideline for information professionals to describe electronic resources and provide access to online informaton for users; it also facilitates the metadata quality control and sharing metadata between different communities and metadata schemes.

Monday, November 10, 2008

Generating MARC with MarcEdit

It has been a trend to harvest metadata from available online resources. Since OAI was adopted by most of data providers, it has facilitated libraries to share metadata . However, sometimes you probably want to integrate a few websites into the library cataloging database, you could do this easily with MarcEdit.

MarcEdit could process the conversion between MARC and XML metadata, it could do the following transformation:
  • MARC→ Dublin Core XML
  • MARC→ MARCXML
Other conversions could be possible, but the above transformations are commonly used by librarians. Users can also edit those marc records with MarcEdit, and batch load them into your intergrating library system. If users could make use of some macros, the bacth editing will be much easier. People who are intertested in this could look at the sample at Miller Library.

Monday, October 6, 2008

Metadata Quality Control

Metadata quality control is becoming more important than ever when we deal with batch import. These problems include various versions of author name, unconventional abbreviations, inappropriate data formats, duplicate records, spelling errors, incomplete punctuations, unrecognized characters etc. To control the quality of metadata, we need systematical way to instantly identify and correct those errors; on the other hand, we also need strengthen metadata creation at the beginning.

Recently, DSpace has released an add-on for metadata quality control, which has powerful mass-edit feature, duplicate detection and resolving algorithms. This is a very promising feature for metadata quality control, at least, some systematical methods could be adopted to solve the current dilemma at the time of submission.

Hopefully, those functions, which are available in traditional Integrating Library System, such as, authority control, control vocabulary etc., will be available in Institutional Repository software.

Sunday, September 28, 2008

Metadata Harvesting with MarcEdit

Last time, when I harvested DSpace metadata with MarcEdit, it only lasted around 3 minutes. From the test on yesterday, it went well even though I only did a small collection.

It is important to ensure consistent metadata and access to those digital collections, especially during the period of senior seminars. A simple and clear processing manual should be written to train staff or student assistants how to create metadata, publish theses and harvest metadata with MarcEdit. I am also interested in harvesting metadata with other software. I hope I can hear some good news from Claude in Montreal soon!

Monday, April 7, 2008

Metadata In DSpace

DSpace uses Dublin Core metadata schema, which includes 15 elements and some qualifiers to adapt to the library implementation profile. DSpace supports OAI (Open Archives Initiative’s Protocol) to provide metadata harvesting as a data provider, which is crucial for Open Archive Initiative.

Dspace could export metadata in DCXML file, now the team is working on migrating the export capability to use the METS standard, which will facilitate to exchange digital library objects between repositories. Basically, DSpace could be customized to allow interoperation with other library or document systems for auto-depositing in DSpace. With this capability, I try to establish the export capability for transferring metadata between dspace and our ILS, then our library users only need one search on library catalog to locate the information in our digital collection, which will definitely improve the search effectiveness and efficiency.

Monday, March 24, 2008

Building A Digital Repository, Part III Metadata Creation

Our first project is Senior Capstone Experience. Metadata creation will start from senior theses. Firstly, we need convert WORD into PDF, which will be read only. We request that all the papers should be submitted with clear title, author full name, optional abstract and keyword, which will be ready to create metadata.

The default metadata scheme on DSpace is Dublin Core. We create and customize metadata according to the feature for each paper. Some papers provides more bibliographic information, such as table of contents, extra contributor to their paper, e.g. advisor, non-control vocabularies for keywords etc., we try to ingest more bibliographic information to our metadata creation, and create more access entries for users. At same time, it also means extra work for us, we need validate subject heading offered by authors. Theses which cross disciplinary will be linked to all related disciplines to ensure users can find them at any related departments. With medatada creation for each record, now users can search them.