Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Access papers from institutional and subject repositories at scale
.claude/skills/brycewang-stanford-institutional-repository-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 92% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 231% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 123% | 0% |
Institutional repositories (IRs) are university-run digital archives that store and provide open access to their researchers' scholarly output — dissertations, journal articles, conference papers, datasets, and technical reports. Subject repositories like arXiv, bioRxiv, SSRN, and RePEc serve similar functions for specific disciplines. Together, they form a distributed network of open scholarship that complements commercial databases.
This guide covers how to discover, access, and systematically harvest content from institutional and subject repositories for literature reviews, meta-analyses, and research data collection.
Institutional Repositories (IR):
- Run by universities to archive their researchers' output
- Examples: DSpace, EPrints, Fedora-based systems
- Discovery: OpenDOAR directory (v2.sherpa.ac.uk/opendoar)
Subject Repositories:
- Discipline-specific archives
- arXiv (physics, CS, math), bioRxiv, SSRN, RePEc, EarthArXiv
Aggregators:
- Harvest from many repositories into a single search interface
- BASE (Bielefeld Academic Search Engine)
- CORE (core.ac.uk, 200M+ open access articles)
- OpenAIRE (European research output)OpenDOAR (Directory of Open Access Repositories) is the primary registry for finding institutional repositories:
pythonimport urllib.request import json def search_opendoar(subject: str = None, country: str = None) -> list: """ Search the OpenDOAR registry for institutional repositories. Args: subject: Filter by subject area (e.g., "Biology", "Computer Science") country: ISO country code (e.g., "US", "GB", "CN") """ base_url = "https://v2.sherpa.ac.uk/cgi/retrieve" params = "?item-type=repository&format=Json" if subject: params += f"&filter=[[\"{subject}\",\"subject\"]]" if country: params += f"&filter=[[\"{country}\",\"country\"]]" req = urllib.request.Request(base_url + params) response = urllib.request.urlopen(req) data = json.loads(response.read()) repositories = [] for item in data.get("items", []): repo_info = { "name": item.get("repository_metadata", {}).get("name", [{}])[0].get("name", ""), "url": item.get("repository_metadata", {}).get("url", ""), "oai_url": item.get("repository_metadata", {}).get("oai_url", ""), "software": item.get("repository_metadata", {}).get("software", {}).get("name", ""), "type": item.get("repository_metadata", {}).get("type", "") } repositories.append(repo_info) return repositories
Most institutional repositories support OAI-PMH (Open Archives Initiative Protocol for Metadata Harvesting), the standard protocol for metadata exchange:
pythonimport xml.etree.ElementTree as ET import urllib.request def harvest_repository(base_url: str, metadata_prefix: str = "oai_dc", set_spec: str = None, from_date: str = None) -> list: """ Harvest metadata records from a repository's OAI-PMH endpoint. Args: base_url: The OAI-PMH base URL metadata_prefix: Metadata format (oai_dc, datacite, mets) set_spec: Optional set/collection to restrict harvesting from_date: Harvest only records added after this date (YYYY-MM-DD) """ params = f"?verb=ListRecords&metadataPrefix={metadata_prefix}" if set_spec: params += f"&set={set_spec}" if from_date: params += f"&from={from_date}" url = base_url + params records = [] while url: response = urllib.request.urlopen(url) tree = ET.parse(response) root = tree.getroot() ns = {"oai": "http://www.openarchives.org/OAI/2.0/"} for record in root.findall(".//oai:record", ns): header = record.find("oai:header", ns) identifier = header.find("oai:identifier", ns).text datestamp = header.find("oai:datestamp", ns).text records.append({"identifier": identifier, "datestamp": datestamp}) token_elem = root.find(".//oai:resumptionToken", ns) if token_elem is not None and token_elem.text: url = f"{base_url}?verb=ListRecords&resumptionToken={token_elem.text}" else: url = None return records
| Verb | Purpose | |------|---------| | Identify | Get repository name, admin email, policies | | ListSets | List available collections/sets | | ListMetadataFormats | List supported metadata schemas | | ListIdentifiers | Lightweight listing of record headers | | ListRecords | Full metadata records with pagination | | GetRecord | Retrieve a single record by identifier |
The most widely deployed open-source repository platform (used by ~40% of repositories worldwide):
{base-url}/oai/request{base-url}/server/apiPopular in the UK and Europe:
{base-url}/cgi/oai2{base-url}/cgi/export/{id}/{format}Used by larger institutions with complex digital collections:
1. Identify target repositories
- Use OpenDOAR to find IRs by subject or country
- List subject repositories relevant to your discipline
2. Test endpoints
- Send Identify request to verify the endpoint is active
- Check ListMetadataFormats for available schemas
3. Harvest incrementally
- Use "from" parameter to harvest only new records
- Store last harvest date for each repository
- Respect rate limits (typically 1 request per second)
4. Deduplicate
- Match records by DOI when available
- Use title + author fuzzy matching for records without DOIs
- Flag duplicates rather than deleting (keep provenance)
5. Store and index
- Save metadata in structured format (JSON, SQLite, CSV)
- Build a local search index for efficient retrievalrobots.txt and repository rate limits| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 20,405 | 36,923 | +81% | 1 | 1 | 0% | 4,099 | 5,531 | +35% | 0 | 0 | — |
case-02 | fail→pass | 21,507 | 25,177 | +17% | 1 | 1 | 0% | 4,366 | 6,266 | +44% | 0 | 0 | — |
case-03 | pass→pass | 12,794 | 13,781 | +8% | 1 | 1 | 0% | 2,203 | 4,240 | +92% | 0 | 0 | — |
case-04 | pass→pass | 5,111 | 6,745 | +32% | 1 | 1 | 0% | 867 | 2,871 | +231% | 0 | 0 | — |
case-05 | pass→pass | 8,153 | 8,464 | +4% | 1 | 1 | 0% | 1,410 | 3,139 | +123% | 0 | 0 | — |
case-06 | pass→pass | 5,870 | 10,283 | +75% | 1 | 1 | 0% | 916 | 3,372 | +268% | 0 | 0 | — |
case-07 | pass→pass | 4,091 | 5,817 | +42% | 1 | 1 | 0% | 617 | 2,877 | +366% | 0 | 0 | — |
case-08 | pass→pass | 3,423 | 4,858 | +42% | 1 | 1 | 0% | 553 | 2,754 | +398% | 0 | 0 | — |
case-09 | pass→pass | 10,201 | 8,807 | -14% | 1 | 1 | 0% | 1,777 | 3,443 | +94% | 0 | 0 | — |
case-10 | pass→pass | 10,777 | 6,461 | -40% | 1 | 1 | 0% | 1,843 | 3,008 | +63% | 0 | 0 | — |
case-11 | pass→pass | 11,419 | 9,294 | -19% | 1 | 1 | 0% | 1,652 | 3,203 | +94% | 0 | 0 | — |
case-12 | pass→pass | 8,262 | 6,330 | -23% | 1 | 1 | 0% | 1,364 | 2,951 | +116% | 0 | 0 | — |
case-18 | pass→pass | 5,414 | 2,712 | -50% | 1 | 1 | 0% | 811 | 2,214 | +173% | 0 | 0 | — |
case-13 | pass→pass | 13,829 | 12,709 | -8% | 1 | 1 | 0% | 2,242 | 4,213 | +88% | 0 | 0 | — |
case-14 | pass→pass | 20,213 | 19,816 | -2% | 1 | 1 | 0% | 2,797 | 4,715 | +69% | 0 | 0 | — |
case-15 | pass→pass | 16,308 | 15,293 | -6% | 1 | 1 | 0% | 2,276 | 4,015 | +76% | 0 | 0 | — |
case-16 | pass→pass | 6,887 | 3,970 | -42% | 1 | 1 | 0% | 976 | 2,488 | +155% | 0 | 0 | — |
case-17 | pass→pass | 9,871 | 5,337 | -46% | 1 | 1 | 0% | 1,475 | 2,751 | +87% | 0 | 0 | — |
case-19 | pass→pass | 6,879 | 2,327 | -66% | 1 | 1 | 0% | 1,033 | 2,238 | +117% | 0 | 0 | — |
case-20 | pass→pass | 6,494 | 5,498 | -15% | 1 | 1 | 0% | 1,184 | 2,939 | +148% | 0 | 0 | — |
case-21 | pass→pass | 17,081 | 18,269 | +7% | 1 | 1 | 0% | 2,477 | 4,641 | +87% | 0 | 0 | — |
case-22 | pass→pass | 14,201 | 17,045 | +20% | 1 | 1 | 0% | 2,617 | 4,824 | +84% | 0 | 0 | — |
case-23 | pass→pass | 16,299 | 17,293 | +6% | 1 | 1 | 0% | 2,681 | 5,039 | +88% | 0 | 0 | — |
case-24 | pass→pass | 11,021 | 11,400 | +3% | 1 | 1 | 0% | 1,915 | 3,892 | +103% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +8 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.