Guides
How a research agent legally gets the full text of a paywalled paper
Given a DOI you cannot read, a handful of free APIs will legally hand you the full text or tell you there is no open copy. Unpaywall tells you where the open copy is and which version it is. Europe PMC gives you parsed XML when the paper is in its open-access subset. CORE finds copies in repositories. OpenAlex and Crossref give you the metadata around the paper, including licence and retraction status. None needs a subscription. When no open copy exists, you get an abstract or nothing. Below are the request for each, the limits as of 2026-09-27, and what Hyperresearch, the AI deep research API, does with them.
Unpaywall finds the open copy and names its version
curl -sS 'https://api.unpaywall.org/v2/10.7717/peerj.4375?email=you@your-domain.org'Read is_oa first, then oa_status, which is one of gold, hybrid, bronze, green or closed. Then read best_oa_location. Its useful fields are url_for_pdf, url_for_landing_page, url, host_type, license and version. The docs define url as “the url_for_pdf if there is one; otherwise landing page URL”. oa_locations lists every copy Unpaywall knows about. You will need it, because the best copy is not always one you can fetch.
For this DOI the call returns is_oa: true, oa_status: "gold", a publishedVersion from the publisher under cc-by, and nine locations. url_for_pdf is null. That is normal. Fall back to url, then try the other locations.
There is no key, but email is required and checked: email=test@example.com returns HTTP 422 with “Please use your own email address in API calls.” The only limit is a request: “Please limit use to 100,000 calls per day.” Two location fields, evidence and updated, are deprecated and always return the string deprecated. Do not fetch from oa_locations_embargoed. The docs describe those copies as “not yet available”.
Europe PMC is the one that returns structured text
The others point you at a PDF. Europe PMC returns JATS XML with real sections and paragraphs, which an agent reads far better than a two-column PDF.
curl -sS 'https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=DOI:%2210.7717/peerj.4375%22&format=json&resultType=core'curl -sS 'https://www.ebi.ac.uk/europepmc/webservices/rest/PMC5815332/fullTextXML'Ask for resultType=core. The default lite result has the PMCID and the open-access flag but leaves out the licence and the full-text links. In the result, check isOpenAccess and inEPMC, and take the PMCID from pmcid. Do not rely on hasFullTextXML: it came back null on this record, which does have XML. In fullTextUrlList.fullTextUrl[], keep only entries with availabilityCode OA. The same record also lists a subscription link.
The second call returns application/xml, about 128 KB, with an <article> root and 88 paragraphs in the body. The same call for an article outside the open-access subset returned HTTP 500 in three tries, so ask only when isOpenAccess is Y.
Open access does not mean anything goes. Europe PMC says articles in the subset “are still protected by copyright” and are made available under a Creative Commons licence “or similar”. Some of those licences are non-commercial. Read license before you reuse text. For bulk work the rule is explicit: “It is not permissible to use any kind of automated process to bulk download other content from Europe PMC.” The open-access subset has its own bulk download. Nothing else does.
There is no API key and no published rate limit. It is not on the developer pages or in the 66-page reference PDF. The only figure is a 2020 post by the Europe PMC team on their own forum: “Currently we apply some throttling of 10 requests per second and 500 requests per minute.” That is six years old. Stay well under it.
CORE casts the widest net, and holds back text without a key
curl -sS 'https://api.core.ac.uk/v3/search/works/?q=doi%3A%2210.7717%2Fpeerj.4375%22'Keep the trailing slash after works. Without it you get a 301. There is no DOI endpoint, so you search on the doi field.
You do not need to register. The docs say “Access to the CORE API is free and requires no authentication”, and the call above works with no key. A key from CORE’s registration form goes in an Authorization: Bearer header. A bad key is worse than none: it returns 401.
The docs’ footnote is the part that matters: “Full-text is not available for unauthenticated API users.” Without a key, fullText comes back as the string "Not available for public API users." An agent that only checks whether the field is empty will store that sentence as the paper. What you do get without a key is downloadUrl (here a CORE-hosted PDF) and sourceFulltextUrls, which for this paper lists the CORE copy and the publisher’s PDF.
Limits are counted in tokens. A simple query costs one, and harder queries cost three to five. Without a key you get “100 tokens per day, maximum 10 per minute”. A registered personal key gets “1,000 tokens per day, maximum 25 per minute”. Registered academic users at a non-supporting institution get “5,000 tokens per day, maximum 10 per minute”. The x-ratelimit-retry-after header is a timestamp, not a number of seconds.
The work record has no version, open-access status or licence field. If you need them, look up the DOI in Unpaywall or OpenAlex.
OpenAlex charges in dollars now, and DOI lookups are free
docs.openalex.org now redirects to help.openalex.org, and many guides still describe the old rules. A free API key, sent as api_key or a Bearer header, is now how you identify yourself.
curl -sS 'https://api.openalex.org/works/doi:10.7717/peerj.4375?api_key=YOUR_KEY'Read open_access and best_oa_location. For this DOI, open_access says is_oa: true, oa_status: "gold" and any_repository_has_fulltext: true. best_oa_location has version: "publishedVersion", license: "cc-by", a null pdf_url, and is_accepted and is_published flags. The record also carries is_retracted and indexed_in. That retraction flag is a reason to call OpenAlex even when you already have the text.
Limits are a daily budget in dollars: $0.10 a day without a key, $1 a day with a free key, reset at midnight UTC. You also get a 429 above 100 requests per second. A single lookup by ID or DOI costs nothing. List and filter calls cost $0.10 per 1,000, search costs $1 per 1,000, and PDF download through their content API costs $10 per 1,000. Their own table lists retrieving 1,000,000 works by DOI as free. The response headers confirm it: X-RateLimit-Cost-USD: 0 on a DOI lookup.
Do not mix up has_fulltext and open_access.any_repository_has_fulltext. On this record the first is false and the second is true. They answer different questions.
Crossref gives the publisher’s licence and text-mining links
curl -sS 'https://api.crossref.org/works/10.7717/peerj.4375?mailto=you@your-domain.org'Crossref does not find open copies, but publishers deposit two useful fields there. license[] holds licence URLs with start dates. link[] holds full-text URLs, and entries marked intended-application: text-mining are where the publisher serves files for machines. A link is not a licence. Check license before you use what it points to.
Add mailto and you move to the polite pool. The response headers show 10 requests per second with 3 at once, against 5 per second with 1 at once without it.
Field names do not line up across services
| What you want | Unpaywall | OpenAlex | Europe PMC | CORE |
|---|---|---|---|---|
| Is it open | is_oa |
open_access.is_oa |
isOpenAccess |
not given |
| Open-access status | oa_status |
open_access.oa_status |
not given | not given |
| Best PDF | best_oa_location.url_for_pdf |
best_oa_location.pdf_url |
fullTextUrlList entries |
downloadUrl |
| Landing page | best_oa_location.url_for_landing_page |
best_oa_location.landing_page_url |
not applicable | not given |
| Every copy | oa_locations[] |
locations[] |
fullTextUrlList[] |
sourceFulltextUrls[] |
| Version | version |
version, is_accepted, is_published |
not given | not given |
| Licence | license |
license, license_id |
license |
not given |
| Retraction flag | not given | is_retracted |
not given | not given |
Record which version you got
Unpaywall uses the DRIVER Guidelines v2.0. submittedVersion “is not yet peer-reviewed”. acceptedVersion “is peer-reviewed, but lacks publisher-specific formatting”. publishedVersion “is the version of record”. They are not the same text. Page numbers differ, and wording and figures can change between the accepted manuscript and the final proof.
Store four things with the text: which service answered, the URL the text came from, the version, and the licence. Then show them. A reader who sees a journal citation should know when the text was an accepted manuscript. A quote from a submitted version may not appear in the published paper.
What not to do
Sci-Hub and its mirrors. Courts in several countries have issued judgments and blocking orders against it. A commercial pipeline cannot use it.
Someone else’s login. Do not use a customer’s or an employee’s institutional login, do not share a proxy session, and do not let an agent sign in anywhere. An institutional licence lets a person read. It does not let a service copy.
Pages behind a paywall. A page that needs payment or a session is not a public page. When a fetch hits one, record it and ask a person.
How Hyperresearch does it
Hyperresearch reads each source through a fetch ladder: the workspace vault, a direct fetch, open access, licensed articles (built but switched off), browser rendering, two managed-fetch rungs, then a person. How a run works covers the ladder. This section covers the open-access step.
The open-access rung runs when the direct fetch was blocked, hit a login wall, failed or came back thin, and the run has an identifier for the paper. For a biomedical article with a DOI, PubMed ID or PMCID, it first makes one Europe PMC lookup and, if the article is in the open-access subset, one full-text fetch. For any DOI, it asks Unpaywall, Europe PMC and CORE where open copies are, then tries them in this order: Unpaywall’s PDF, Unpaywall’s landing page, Europe PMC, then a CORE repository copy. It fetches at most three copies for one source, and it doesn’t look an article up on Europe PMC twice. The recovered text must pass the same length and junk checks as any other page. If no copy passes, the ladder moves on.
Every URL a resolver hands back is checked before anything fetches it, because it arrives inside someone else’s JSON. The check covers the scheme, credentials in the URL, and whether the host resolves to a public address, and it runs again on every redirect.
You can call the same lookup yourself. It needs the vault:read scope and returns every location it found, each with its version, licence and whether it passed the URL check.
curl -sS https://api.hyperresearch.ai/v1/scholar/oa \ -H 'Authorization: Bearer <YOUR_API_KEY>' \ -H 'Content-Type: application/json' \ -d '{"doi": "10.7717/peerj.4375"}'A recovered source is labelled one of two ways. substituted means the original page loaded but was thin, so only the body came from the open copy. rescued means nothing came from the original page, so the title and authors are the open copy’s too. Each source row in the run result has an oa block with the kind, the resolver, the URL the text came from, the version and body_is_not_from_source: true. A rescued source also has nothing_from_source: true. When the resolver reported a licence, the note keeps it. The source is kept in the vault with the rung that read it.
The standalone citation verification API reads cited pages through a shorter ladder. Its open-access step uses only the Europe PMC path for biomedical articles.
We have not measured how often each rung recovers a paper, so this page gives no recovery rate.
The rules are the ones the open-source harness keeps, described in Claude Code as a research harness: public pages, logged out, no credentials, no CAPTCHA solving. When a paper needs a person, the run gives you a task. You can upload a copy you hold.
By Jordan Gibbs · Updated 2026-09-27
