Building a lightweight dataset catalog in Digital Commons

Document Type

Article

Publication Title

Journal of eScience Librarianship

Abstract

Introduction: Many academic libraries compile dataset catalogs to make the research data produced by their institutions more findable, accessible, interoperable, and reusable (FAIR). There are a variety of approaches to gather metadata for dataset records and catalogs may be hosted on one of several platforms.

Methods: We developed Python code and leveraged APIs to locate relevant datasets and harvest their metadata automatically. The metadata is then manually curated and enhanced. We employ the hosted institutional repository Digital Commons to display the dataset records.

Results: The University of Alabama at Birmingham’s Research Data Catalog (RDC) currently contains over 260 dataset records from multiple generalist repositories. The code we developed, and the customized Digital Commons collection, are available for reuse.

Conclusion: Combining API harvesting with the ready-made features of Digital Commons yielded efficient ingestion of many dataset records, allowing us to prioritize manual curation and enhancement of the dataset metadata. This approach is ideal for launching a dataset catalog at institutions with limited personnel time and minimal technical resources.

DOI

https://doi.org/10.7191/jeslib.1159

Publication Date

12-16-2025

College or School

UAB Libraries

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.

Share

COinS