Rank CS/AI/ML papers by citations, forecast impact, or code adoption; trace 23.2M citation edges.
io.github.YGao2005/scholar-feed-mcp: Scholar Feed MCP Server
This MCP server ranks CS/AI/ML papers by citations, forecasted impact, or code adoption, and supports tracing citation relationships (23.2M citation edges). It is associated with topics including arXiv, semantic search, BibTeX, and developer workflows involving Claude and Cursor.
π οΈ Key Features
Rank papers by citations
Forecast impact
Rank by code adoption
Trace citation edges (23.2M)
π Use Cases
Literature review over CS/AI/ML papers
Semantic search within research corpora
Building paper discovery workflows that include BibTeX outputs
Supporting Claude/Claude Code and Cursor-based research
β‘ Developer Benefits
Provides paper ranking and citation tracing signals
Works within the Model Context Protocol (MCP) ecosystem
Search Scholar Feed's 600k+ CS/AI/ML paper corpus. Semantic (embedding) search by default, so it finds conceptually related work even when the wording differs. EVERY PARAMETER DOCUMENTS ITS OWN BEHAVIOUR AND COVERAGE LIMITS β read the ones you intend to use; this description covers only what no single parameter can tell you. RETRIEVAL LIMIT: semantic ranking favours recent, stylistically-matched papers and routinely MISSES the old high-citation anchor of a field (H2O for KV eviction, GRIT for unified embedding+generation). To reach a field's canonical work, read the top-5 abstracts for repeated baseline mentions ('we compare against X') and look that name up directly, or call get_foundational_lineage. THREE UNRELATED NOTIONS OF IMPACT, easily confused: proven citations (sort='impactful', min_citations) | a ~90-day forecast percentile that is NULL on older papers and therefore excludes them (sort='trending', impact_min) | GitHub adoption (sort='community', min_stars). YOUR LIBRARY IS MARKED INLINE on authenticated calls: each hit carries is_saved and is_read, and a hit you previously annotated carries note_text β your own earlier verdict. Read note_text INSTEAD of re-deriving a conclusion from the abstract; re-judging a paper you already ruled on is the most common way an agent wastes a research session. is_saved=false is a real measurement; on anonymous calls these keys are absent entirely, so never read a missing is_saved as false. Papers new to you are ranked exactly as before β nothing is demoted for being unseen.
Parameters27
q
string
optional
Search query keywords. REQUIRED unless (a) anchor_paper_id or scope_to_citations_of is set (anchor mode ignores q and returns papers similar to the anchor), or (b) sort is one of 'trending' / 'recent' / 'impactful' β the query-less 'browse the frontier' feed, which is capped to the FIRST 200 RESULTS (paging past offset 200 is a 422; narrow with q= or filters instead). Any other q-less call is a 422, INCLUDING a q-less call with only filters (category/days/...) and a q-less sort='community'. Filters alone do NOT substitute for q β pair them with a browse sort (e.g. category='cs.AI' + sort='recent') or pass q. For 'what's hot in AI right now', either sort='trending' alone or a broad q plus sort='trending' works.
sort
string
optional
Result ranking β a relevanceβimpact dial plus time-based and adoption orders. 'relevance' (default) = best topical match. 'balanced' = relevant AND well-cited. 'impactful' = the most-cited (proven-influential) papers among those relevant to the query β use this for 'the important/seminal papers on topic X'. 'trending' = rising/FORECAST impact (impact_pct, last ~90 days) β use for 'what's hot/new in X', NOT for established work. 'recent' = newest first. 'community' = GitHub adoption (stars + star-velocity) β surfaces the papers practitioners are actually running/building on, independent of citations. COVERAGE CAVEAT (the analogue of impact_min's ~90-day hole): an unfetched repo stores 0 rather than NULL, and coverage skews heavily toward recently-published papers, so most older papers with a repo currently rank as 0-star and sink β 'community' reflects measured adoption, not corpus-wide adoption. Proven impact ('impactful'/'balanced') ranks by real citations; 'trending' is a model prediction; 'community' is real-world engineering traction within its window. Pair with get_foundational_lineage for a topic's canonical roots.
anchor_paper_id
string
optional
Return papers similar to this arXiv paper ID. When set, q is ignored and results carry similarity_score. Example: '2407.15831'.
scope_to_citations_of
string
optional
Restrict search to this paper's citation graph, ranked by relevance to q. Pass the arXiv ID of the paper whose citations you want to search within.
category
string
optional
Filter by arXiv category e.g. 'cs.AI', 'cs.LG'
novelty_min
number
optional
Minimum novelty score (0-1). Use 0.5+ for novel papers.
impact_min
integer
optional
Minimum impact_pct (0-100), e.g. 80 = top 20% FORECAST impact. This is a RISING-WORK filter: impact_pct is only computed for the last ~90 days, so impact_min restricts results to recent papers predicted to land well AND DROPS everything older. Use it for 'what's rising in X'. Do NOT use it to find the influential/seminal papers in a topic β that excludes the established work; use sort='impactful' instead.
days
integer
optional
Limit to papers published within N days
has_code
boolean
optional
Filter to papers with a linked code release (has_code=true). Surfaces runnable/reproducible work β pair with min_stars/sort='community' to find the papers practitioners actually adopt.
min_citations
integer
optional
Minimum real citation count. Unlike impact_min (a ~90-day FORECAST percentile), this filters on PROVEN citations and keeps established/canonical papers.
min_stars
integer
optional
Minimum GitHub stars on the paper's linked repo. A proxy for engineering adoption β surfaces work that practitioners are actually running/building on. Pair with sort='community' to rank by it. COVERAGE CAVEAT: a never-fetched repo is stored as 0, not NULL, so this filter cannot distinguish 'no adoption' from 'never measured'. It is applied as 'KNOWN to have >= N stars' β papers whose stars were never fetched are excluded rather than treated as 0-star, so the result is honest but INCOMPLETE: a genuinely popular older paper can be missing simply because nobody measured it. Coverage skews toward recently-published papers and is being backfilled. Use it to filter recent work; for established papers use min_citations instead.
github_url_exists
boolean
optional
Filter on whether the paper has a linked GitHub URL (true = only papers with a repo). Stricter than has_code (which counts any code link).
published_after
string
optional
Only papers published on or after this date, 'YYYY-MM-DD'. Use with published_before to bound an arbitrary date window (days only gives a rolling N-day lookback).
published_before
string
optional
Only papers published on or before this date, 'YYYY-MM-DD'. Pair with published_after for an explicit window.
method_category
string
optional
Filter by method category e.g. 'reinforcement learning', 'transformer'
method_name
string
optional
Filter to papers introducing/using a specific named method e.g. 'LoRA', 'YOLO', 'DPO'. Case-insensitive substring match on the extracted method_name field.
task
string
optional
Filter by task e.g. 'image classification', 'question answering' (partial match)
dataset
string
optional
Filter to papers that evaluate on a specific dataset e.g. 'MMLU', 'ImageNet'
contribution_type
string
optional
Filter by paper's contribution type
task_category
string
optional
Filter by broad research area
mode
string
optional
Search mode. 'semantic' (default) uses embedding similarity β finds conceptually related papers even without exact keyword matches. 'keyword' uses Postgres full-text search β faster but only matches exact terms.
cursor
string
optional
Cursor from previous response's next_cursor for keyset pagination
page
integer
optional
Page number
limit
integer
optional
Results per page (max 50)
fields
string
optional
Comma-separated list of fields to return (e.g. 'arxiv_id,title,llm_summary,llm_novelty_score'). If omitted, returns the lean 12-field default unless verbose=true.
verbose
boolean
optional
If true, returns the full 28-field paper shape (method/task/dataset extraction, application_domain, baselines, etc.). Default false returns the lean 12-field set. Ignored when `fields` is provided.
exclude_ids
array
optional
arXiv IDs to exclude from results (for deduplication across chained calls)
Raw schema
{
"type": "object",
"properties": {
"q": {
"description": "Search query keywords. REQUIRED unless (a) anchor_paper_id or scope_to_citations_of is set (anchor mode ignores q and returns papers similar to the anchor), or (b) sort is one of 'trending' / 'recent' / 'impactful' β the query-less 'browse the frontier' feed, which is capped to the FIRST 200 RESULTS (paging past offset 200 is a 422; narrow with q= or filters instead). Any other q-less call is a 422, INCLUDING a q-less call with only filters (category/days/...) and a q-less sort='community'. Filters alone do NOT substitute for q β pair them with a browse sort (e.g. category='cs.AI' + sort='recent') or pass q. For 'what's hot in AI right now', either sort='trending' alone or a broad q plus sort='trending' works.",
"type": "string",
"minLength": 1
},
"sort": {
"description": "Result ranking β a relevanceβimpact dial plus time-based and adoption orders. 'relevance' (default) = best topical match. 'balanced' = relevant AND well-cited. 'impactful' = the most-cited (proven-influential) papers among those relevant to the query β use this for 'the important/seminal papers on topic X'. 'trending' = rising/FORECAST impact (impact_pct, last ~90 days) β use for 'what's hot/new in X', NOT for established work. 'recent' = newest first. 'community' = GitHub adoption (stars + star-velocity) β surfaces the papers practitioners are actually running/building on, independent of citations. COVERAGE CAVEAT (the analogue of impact_min's ~90-day hole): an unfetched repo stores 0 rather than NULL, and coverage skews heavily toward recently-published papers, so most older papers with a repo currently rank as 0-star and sink β 'community' reflects measured adoption, not corpus-wide adoption. Proven impact ('impactful'/'balanced') ranks by real citations; 'trending' is a model prediction; 'community' is real-world engineering traction within its window. Pair with get_foundational_lineage for a topic's canonical roots.",
"type": "string",
"enum": [
"relevance",
"balanced",
"impactful",
"trending",
"recent",
"community"
]
},
"anchor_paper_id": {
"description": "Return papers similar to this arXiv paper ID. When set, q is ignored and results carry similarity_score. Example: '2407.15831'.",
"type": "string"
},
"scope_to_citations_of": {
"description": "Restrict search to this paper's citation graph, ranked by relevance to q. Pass the arXiv ID of the paper whose citations you want to search within.",
"type": "string"
},
"category": {
"description": "Filter by arXiv category e.g. 'cs.AI', 'cs.LG'",
"type": "string"
},
"novelty_min": {
"description": "Minimum novelty score (0-1). Use 0.5+ for novel papers.",
"type": "number",
"minimum": 0,
"maximum": 1
},
"impact_min": {
"description": "Minimum impact_pct (0-100), e.g. 80 = top 20% FORECAST impact. This is a RISING-WORK filter: impact_pct is only computed for the last ~90 days, so impact_min restricts results to recent papers predicted to land well AND DROPS everything older. Use it for 'what's rising in X'. Do NOT use it to find the influential/seminal papers in a topic β that excludes the established work; use sort='impactful' instead.",
"type": "integer",
"minimum": 0,
"maximum": 100
},
"days": {
"description": "Limit to papers published within N days",
"type": "integer",
"minimum": 1,
"maximum": 3650
},
"has_code": {
"description": "Filter to papers with a linked code release (has_code=true). Surfaces runnable/reproducible work β pair with min_stars/sort='community' to find the papers practitioners actually adopt.",
"type": "boolean"
},
"min_citations": {
"description": "Minimum real citation count. Unlike impact_min (a ~90-day FORECAST percentile), this filters on PROVEN citations and keeps established/canonical papers.",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"min_stars": {
"description": "Minimum GitHub stars on the paper's linked repo. A proxy for engineering adoption β surfaces work that practitioners are actually running/building on. Pair with sort='community' to rank by it. COVERAGE CAVEAT: a never-fetched repo is stored as 0, not NULL, so this filter cannot distinguish 'no adoption' from 'never measured'. It is applied as 'KNOWN to have >= N stars' β papers whose stars were never fetched are excluded rather than treated as 0-star, so the result is honest but INCOMPLETE: a genuinely popular older paper can be missing simply because nobody measured it. Coverage skews toward recently-published papers and is being backfilled. Use it to filter recent work; for established papers use min_citations instead.",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"github_url_exists": {
"description": "Filter on whether the paper has a linked GitHub URL (true = only papers with a repo). Stricter than has_code (which counts any code link).",
"type": "boolean"
},
"published_after": {
"description": "Only papers published on or after this date, 'YYYY-MM-DD'. Use with published_before to bound an arbitrary date window (days only gives a rolling N-day lookback).",
"type": "string"
},
"published_before": {
"description": "Only papers published on or before this date, 'YYYY-MM-DD'. Pair with published_after for an explicit window.",
"type": "string"
},
"method_category": {
"description": "Filter by method category e.g. 'reinforcement learning', 'transformer'",
"type": "string"
},
"method_name": {
"description": "Filter to papers introducing/using a specific named method e.g. 'LoRA', 'YOLO', 'DPO'. Case-insensitive substring match on the extracted method_name field.",
"type": "string"
},
"task": {
"description": "Filter by task e.g. 'image classification', 'question answering' (partial match)",
"type": "string"
},
"dataset": {
"description": "Filter to papers that evaluate on a specific dataset e.g. 'MMLU', 'ImageNet'",
"type": "string"
},
"contribution_type": {
"description": "Filter by paper's contribution type",
"type": "string",
"enum": [
"model",
"method",
"benchmark",
"dataset",
"survey",
"theoretical",
"empirical_study",
"system"
]
},
"task_category": {
"description": "Filter by broad research area",
"type": "string",
"enum": [
"NLP",
"Computer Vision",
"RL",
"Audio/Speech",
"Graphs",
"Multimodal",
"Systems",
"Theory",
"Security",
"Other"
]
},
"mode": {
"description": "Search mode. 'semantic' (default) uses embedding similarity β finds conceptually related papers even without exact keyword matches. 'keyword' uses Postgres full-text search β faster but only matches exact terms.",
"type": "string",
"enum": [
"keyword",
"semantic"
]
},
"cursor": {
"description": "Cursor from previous response's next_cursor for keyset pagination",
"type": "string"
},
"page": {
"default": 1,
"description": "Page number",
"type": "integer",
"minimum": 1,
"maximum": 9007199254740991
},
"limit": {
"default": 20,
"description": "Results per page (max 50)",
"type": "integer",
"minimum": 1,
"maximum": 50
},
"fields": {
"description": "Comma-separated list of fields to return (e.g. 'arxiv_id,title,llm_summary,llm_novelty_score'). If omitted, returns the lean 12-field default unless verbose=true.",
"type": "string"
},
"verbose": {
"description": "If true, returns the full 28-field paper shape (method/task/dataset extraction, application_domain, baselines, etc.). Default false returns the lean 12-field set. Ignored when `fields` is provided.",
"type": "boolean"
},
"exclude_ids": {
"description": "arXiv IDs to exclude from results (for deduplication across chained calls)",
"type": "array",
"items": {
"type": "string"
}
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}
get_paper
Get full details for one or more papers by arXiv ID. Pass a single-element array for one paper; pass multiple IDs to batch-fetch up to 50 papers in one call. Pass format='bibtex' to get a .bib citation entry (bibtex is single-paper only; for multi-paper bibtex, call repeatedly). Default returns a lean 13-field shape (arxiv_id, title, authors, year, categories, has_code, github_url, citation_count, venue_name, llm_summary, llm_significance, llm_novelty_score, impact_pct β where impact_pct is the ML-forecast impact percentile 0-100 computed WITHIN the paper's own arXiv-category cohort, so it is a cohort-relative rank rather than an absolute score, and is NULL on older papers outside the recent ~90-day scoring window). Pass verbose=true for the full shape with structured extraction (method_name, contribution_type, task_category, datasets, baselines) and institution_tags. Use fields='arxiv_id,title,abstract' to select an exact subset, or fetch_fulltext with sections='all' for the full paper.
Parameters4
arxiv_ids
array
required
One or more arXiv IDs. Single-paper lookup uses [id]; batch lookup passes multiple IDs (max 50). Example: ['2407.15831'] or ['2407.15831', '2402.09906'].
format
string
optional
Response format. 'json' (default) returns structured paper data. 'bibtex' returns a .bib citation entry. Bibtex mode uses the first ID in arxiv_ids.
fields
string
optional
Comma-separated list of fields to return (e.g. 'arxiv_id,title,llm_summary,abstract'). If omitted, returns the lean 12-field default unless verbose=true.
verbose
boolean
optional
If true, returns the full 28-field paper shape (method/task/dataset extraction, application_domain, baselines, etc.). Default false returns the lean 12-field set. Ignored when `fields` is provided.
Raw schema
{
"type": "object",
"properties": {
"arxiv_ids": {
"minItems": 1,
"maxItems": 50,
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"description": "One or more arXiv IDs. Single-paper lookup uses [id]; batch lookup passes multiple IDs (max 50). Example: ['2407.15831'] or ['2407.15831', '2402.09906']."
},
"format": {
"description": "Response format. 'json' (default) returns structured paper data. 'bibtex' returns a .bib citation entry. Bibtex mode uses the first ID in arxiv_ids.",
"type": "string",
"enum": [
"json",
"bibtex"
]
},
"fields": {
"description": "Comma-separated list of fields to return (e.g. 'arxiv_id,title,llm_summary,abstract'). If omitted, returns the lean 12-field default unless verbose=true.",
"type": "string"
},
"verbose": {
"description": "If true, returns the full 28-field paper shape (method/task/dataset extraction, application_domain, baselines, etc.). Default false returns the lean 12-field set. Ignored when `fields` is provided.",
"type": "boolean"
}
},
"required": [
"arxiv_ids"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
get_citations
Get the citation graph for a paper, sorted by citing-paper rank_score (highest-impact first). 'citing' = outgoing references this paper cites; 'cited_by' = incoming citations from other papers. Default response is a lean 12-field shape per paper β pass verbose=true for the full 28-field shape.
Parameters6
arxiv_id
string
required
arXiv ID of the paper
direction
string
optional
'citing' = outgoing references this paper cites; 'cited_by' = incoming citations from other papers
limit
integer
optional
Number of papers to return (max 50)
fields
string
optional
Comma-separated list of fields to return (e.g. 'arxiv_id,title,llm_summary,llm_novelty_score'). If omitted, returns the lean 12-field default unless verbose=true.
verbose
boolean
optional
If true, returns the full 28-field paper shape. Default false returns the lean 12-field set. Ignored when `fields` is provided.
exclude_ids
array
optional
arXiv IDs to exclude from results (for deduplication across chained calls)
Raw schema
{
"type": "object",
"properties": {
"arxiv_id": {
"type": "string",
"minLength": 1,
"description": "arXiv ID of the paper"
},
"direction": {
"default": "cited_by",
"description": "'citing' = outgoing references this paper cites; 'cited_by' = incoming citations from other papers",
"type": "string",
"enum": [
"citing",
"cited_by"
]
},
"limit": {
"default": 20,
"description": "Number of papers to return (max 50)",
"type": "integer",
"minimum": 1,
"maximum": 50
},
"fields": {
"description": "Comma-separated list of fields to return (e.g. 'arxiv_id,title,llm_summary,llm_novelty_score'). If omitted, returns the lean 12-field default unless verbose=true.",
"type": "string"
},
"verbose": {
"description": "If true, returns the full 28-field paper shape. Default false returns the lean 12-field set. Ignored when `fields` is provided.",
"type": "boolean"
},
"exclude_ids": {
"description": "arXiv IDs to exclude from results (for deduplication across chained calls)",
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"arxiv_id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
fetch_fulltext
Extract paper content from an arXiv paper's LaTeX source, falling back to PDF text. Two modes: 'results' (default) returns ~800 chars of results/experiments + up to 3 table captions β lean, ideal for checking a reported number. 'all' returns full paper sections (abstract, introduction, related work, method, results, conclusion) at up to 3000 chars each + 5 table captions, ~15KB, so prefer 'results' unless you need the whole paper. Content is available for ~95% of arXiv papers; a 404 means neither LaTeX nor PDF extraction yielded text. May take a few seconds.
Parameters2
arxiv_id
string
required
arXiv ID of the paper
sections
string
optional
'results' (default): lean ~800-char results/experiments excerpt + table captions. 'all': full paper (abstract, intro, method, results, conclusion, related work) β much larger payload.
Raw schema
{
"type": "object",
"properties": {
"arxiv_id": {
"type": "string",
"minLength": 1,
"description": "arXiv ID of the paper"
},
"sections": {
"description": "'results' (default): lean ~800-char results/experiments excerpt + table captions. 'all': full paper (abstract, intro, method, results, conclusion, related work) β much larger payload.",
"type": "string",
"enum": [
"results",
"all"
]
}
},
"required": [
"arxiv_id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
find_author
Two-mode author tool. Provide exactly one of q or id. Q-MODE (q=...): search for researchers by topic or name β uses embedding similarity for topics ('efficient LLM inference'), fuzzy matching for names ('Yann LeCun'). Returns a list of matching authors with author_id, name, h_index, total_papers, primary_field, research_topics. ID-MODE (id=...): look up a single author profile by author_id (obtained from a previous q-mode call or from co_author_graph results). Returns h-index, total citations, global rank, primary field, novelty score distribution, research topics, code/venue scores, years active, and their top 10 papers by rank score.
Parameters4
q
string
optional
Topic or researcher name to search (q-mode). Returns a list of matching authors. Examples: 'efficient transformer training', 'Geoffrey Hinton'.
id
integer
optional
Author ID for direct profile lookup (id-mode). Returns the single author profile with top 10 papers. Get IDs from q-mode results or co_author_graph.
field
string
optional
(q-mode only) Filter by primary research field e.g. 'cs.LG', 'cs.CV', 'cs.CL'.
limit
integer
optional
(q-mode only) Max results to return (default 20).
Raw schema
{
"type": "object",
"properties": {
"q": {
"description": "Topic or researcher name to search (q-mode). Returns a list of matching authors. Examples: 'efficient transformer training', 'Geoffrey Hinton'.",
"type": "string",
"minLength": 2
},
"id": {
"description": "Author ID for direct profile lookup (id-mode). Returns the single author profile with top 10 papers. Get IDs from q-mode results or co_author_graph.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"field": {
"description": "(q-mode only) Filter by primary research field e.g. 'cs.LG', 'cs.CV', 'cs.CL'.",
"type": "string"
},
"limit": {
"default": 20,
"description": "(q-mode only) Max results to return (default 20).",
"type": "integer",
"minimum": 1,
"maximum": 50
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}
co_author_graph
Find the co-authorship neighborhood of one or more authors. Given a list of author_ids, returns edges {from, to, papers_count, last_collab_year} where 'from' is one of the input authors and 'to' is any co-author appearing on a shared paper within the window. Use for AC reviewer triage (find conflicts), disambiguating researchers (who do they actually work with?), or expanding an author seed into a research community. window_years defaults to 10. Result is capped at 500 edges, sorted by papers_count DESC.
Parameters2
author_ids
array
required
Author IDs to query (1-25). Get author IDs via the find_author tool.
window_years
integer
optional
Only count co-authorships from the last N years (default 10, max 30).
Raw schema
{
"type": "object",
"properties": {
"author_ids": {
"minItems": 1,
"maxItems": 25,
"type": "array",
"items": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"description": "Author IDs to query (1-25). Get author IDs via the find_author tool."
},
"window_years": {
"default": 10,
"description": "Only count co-authorships from the last N years (default 10, max 30).",
"type": "integer",
"minimum": 1,
"maximum": 30
}
},
"required": [
"author_ids"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
embed_text
Embed a text string into a 768-dim Gemini Flash vector. Use for HyDE-style retrieval: (1) write a hypothetical short paper that would perfectly answer the user's query, (2) embed it with task_type='RETRIEVAL_DOCUMENT' (default β matches the corpus embedding side), (3) pass the resulting embedding back through search-style tools to find real papers nearest to the hypothetical. task_type='RETRIEVAL_QUERY' matches the query side and is useful for direct user-query embedding without HyDE. Pro-only β requires an SF_API_KEY on a Pro account; anonymous and free callers get a 403 pro_required. Cost: ~$0.0001/call; rate-limited at 30/minute per API key.
Parameters2
text
string
required
Text to embed (1-8000 chars). For HyDE flows this is your hypothetical answer/abstract.
task_type
string
optional
RETRIEVAL_DOCUMENT (default) matches paper-side embeddings β use for HyDE. RETRIEVAL_QUERY matches query-side semantic search.
Raw schema
{
"type": "object",
"properties": {
"text": {
"type": "string",
"minLength": 1,
"maxLength": 8000,
"description": "Text to embed (1-8000 chars). For HyDE flows this is your hypothetical answer/abstract."
},
"task_type": {
"default": "RETRIEVAL_DOCUMENT",
"description": "RETRIEVAL_DOCUMENT (default) matches paper-side embeddings β use for HyDE. RETRIEVAL_QUERY matches query-side semantic search.",
"type": "string",
"enum": [
"RETRIEVAL_DOCUMENT",
"RETRIEVAL_QUERY"
]
}
},
"required": [
"text"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
get_field_orientation
Returns CANDIDATE FOUNDATIONAL PAPERS for a research topic β cheap retrieval only, no synthesis. Ranks papers by a blend of citation count (0.6 weight, captures importance) and semantic similarity to your topic (0.4 weight). Use this to bootstrap a literature survey or get a fast sense of the landscape. For a synthesized orientation report (key concepts, open problems, reading order), use the /field-guide skill which calls this tool internally. Does not require a Pro API key β no LLM calls are made.
Parameters2
topic
string
required
Research area to orient on. Be specific for better results. Examples: 'diffusion models for protein structure prediction', 'efficient attention mechanisms for long-context LLMs', 'graph neural networks for molecular property prediction'.
limit
integer
optional
Number of candidate papers to return (5β30, default 15).
Raw schema
{
"type": "object",
"properties": {
"topic": {
"type": "string",
"minLength": 5,
"maxLength": 300,
"description": "Research area to orient on. Be specific for better results. Examples: 'diffusion models for protein structure prediction', 'efficient attention mechanisms for long-context LLMs', 'graph neural networks for molecular property prediction'."
},
"limit": {
"default": 15,
"description": "Number of candidate papers to return (5β30, default 15).",
"type": "integer",
"minimum": 5,
"maximum": 30
}
},
"required": [
"topic"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
get_foundational_lineage
Returns the FOUNDATIONAL WORK FOR A PAPER'S NICHE via the citation graph β the relative question ('what is foundational for THIS paper's specific sub-field', often itself only modestly cited) rather than the obvious global landmarks. Anchors on the paper, takes its embedding neighbourhood as the niche, and ranks what the niche cites into three tiers: `niche_roots` (the niche-specific foundations, ranked by how specifically the neighbourhood builds on them β surfaces canonical anchors that semantic search misses), `field_level` (broader secondary foundations), and `discipline` (universal landmarks like Attention Is All You Need, collapsed out of the way). Each paper carries `cited_by_in_niche` evidence so the claim is grounded, not asserted. Use this to trace prior art / lineage for a paper, or to find the canonical methods a niche is built on. Complements get_field_orientation (which is topic-anchored and retrieval-only). No Pro key and no LLM calls required.
Parameters4
anchor_paper_id
string
required
arXiv ID of the paper to anchor on, e.g. '2504.04704' or '2504.04704v2'. The niche is built from this paper's embedding neighbourhood.
When true (default), demote universally-cited landmark papers into the collapsed `discipline` tier so the niche-specific foundations lead. Set false to keep landmarks in the foundational tiers.
limit
integer
optional
Max papers in each of the niche_roots and field_level tiers (5β40, default 15).
Raw schema
{
"type": "object",
"properties": {
"anchor_paper_id": {
"type": "string",
"minLength": 4,
"maxLength": 40,
"description": "arXiv ID of the paper to anchor on, e.g. '2504.04704' or '2504.04704v2'. The niche is built from this paper's embedding neighbourhood."
},
"scope": {
"default": "field",
"description": "Niche breadth: 'narrow' (~100 nearest papers, tightest sub-topic β surfaces the few-citation niche root), 'field' (~200, default), 'broad' (~400, wider area foundations).",
"type": "string",
"enum": [
"narrow",
"field",
"broad"
]
},
"generality_ceiling": {
"default": true,
"description": "When true (default), demote universally-cited landmark papers into the collapsed `discipline` tier so the niche-specific foundations lead. Set false to keep landmarks in the foundational tiers.",
"type": "boolean"
},
"limit": {
"default": 15,
"description": "Max papers in each of the niche_roots and field_level tiers (5β40, default 15).",
"type": "integer",
"minimum": 5,
"maximum": 40
}
},
"required": [
"anchor_paper_id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
save_paper
Save a paper to the authenticated user's Scholar Feed library (bookmark). MUTATES the library and feeds the user's personalization β saved papers are the strongest signal in the For You feed and the email digest. Idempotent: calling it again on an already-saved paper leaves it saved. Requires SF_API_KEY. To file it into a named collection in one step, use add_to_collection (that also saves).
Parameters1
arxiv_id
string
required
arXiv ID of the paper to save, e.g. '2407.15831'.
Raw schema
{
"type": "object",
"properties": {
"arxiv_id": {
"type": "string",
"minLength": 1,
"description": "arXiv ID of the paper to save, e.g. '2407.15831'."
}
},
"required": [
"arxiv_id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
unsave_paper
Remove a paper from the authenticated user's Scholar Feed library. MUTATES the library. Idempotent: removing a paper that isn't saved leaves it unsaved. Note: the saved library is a superset of all collections, so un-saving a paper ALSO removes it from every collection it was in. To keep it filed in a collection, use remove_from_collection instead (that leaves the paper saved). Requires SF_API_KEY.
Parameters1
arxiv_id
string
required
arXiv ID of the paper to remove from the library.
Raw schema
{
"type": "object",
"properties": {
"arxiv_id": {
"type": "string",
"minLength": 1,
"description": "arXiv ID of the paper to remove from the library."
}
},
"required": [
"arxiv_id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
like_paper
Like a paper β a 'more like this' calibration signal that tunes the user's For You feed toward similar work. INSERT-only and idempotent (liking twice is a no-op, never un-likes). Distinct from save_paper: like expresses taste for ranking; save bookmarks for later reading. Requires SF_API_KEY.
Parameters1
arxiv_id
string
required
arXiv ID of the paper to like.
Raw schema
{
"type": "object",
"properties": {
"arxiv_id": {
"type": "string",
"minLength": 1,
"description": "arXiv ID of the paper to like."
}
},
"required": [
"arxiv_id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
list_library
List the authenticated user's saved papers (their library), newest first. Read-only. Use this to review a reading list or to see what's already saved before saving more. Requires SF_API_KEY. SHAPE: agent callers get a lean record β `llm_summary` (~300 chars) INSTEAD of the abstract, with empty fields omitted rather than sent as null. Each paper also carries the state that makes this a knowledge base rather than a bookmark list: `note_text` (the user's own recorded verdict, when one exists), `is_read`, and `collections` (the axes it is filed under, e.g. 'AgentOPA/G4'). READ note_text FIRST. A paper carrying one was already judged in an earlier session β use that verdict instead of re-reading the paper and re-deriving it. If it is missing, consider recording one with annotate_paper so the next session inherits your conclusion. Pass verbose=true (or fields=...) only when you genuinely need the abstract or the full 28-field shape; the default is ~4x smaller and is the right choice for surveying what you already have.
Parameters4
limit
integer
optional
How many saved papers to return (max 100).
page
integer
optional
Page number for paging through a large library.
verbose
boolean
optional
Return the full paper shape (including the abstract) instead of the lean default. Costs roughly 4x the tokens β prefer llm_summary unless you specifically need the abstract's wording.
fields
string
optional
Comma-separated fields to return, e.g. 'arxiv_id,title,abstract'. Overrides verbose. Library state (note_text/is_read/is_saved/collections) is always included regardless.
Raw schema
{
"type": "object",
"properties": {
"limit": {
"default": 50,
"description": "How many saved papers to return (max 100).",
"type": "integer",
"minimum": 1,
"maximum": 100
},
"page": {
"default": 1,
"description": "Page number for paging through a large library.",
"type": "integer",
"minimum": 1,
"maximum": 9007199254740991
},
"verbose": {
"description": "Return the full paper shape (including the abstract) instead of the lean default. Costs roughly 4x the tokens β prefer llm_summary unless you specifically need the abstract's wording.",
"type": "boolean"
},
"fields": {
"description": "Comma-separated fields to return, e.g. 'arxiv_id,title,abstract'. Overrides verbose. Library state (note_text/is_read/is_saved/collections) is always included regardless.",
"type": "string"
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}
list_collections
List the authenticated user's collections (named groups of saved papers) with paper counts. Read-only. Use before add_to_collection to see existing collections. Requires SF_API_KEY.
Create a new named collection. MUTATES. If a collection with that name already exists, returns the existing one (get-or-create β never errors on duplicate). Use "/" to nest under a folder, e.g. "AgentOPA/Formal" β the folder is derived from the name, so there is no parent to create first. Requires SF_API_KEY.
Parameters1
name
string
required
Name for the collection, e.g. 'KV-cache compression'. Use '/' to nest: 'AgentOPA/Formal' files it under an 'AgentOPA' folder.
Raw schema
{
"type": "object",
"properties": {
"name": {
"type": "string",
"minLength": 1,
"maxLength": 100,
"description": "Name for the collection, e.g. 'KV-cache compression'. Use '/' to nest: 'AgentOPA/Formal' files it under an 'AgentOPA' folder."
}
},
"required": [
"name"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
add_to_collection
Add a paper to a collection, addressed by collection_id OR collection_name (get-or-create by name β no need to look up an id first). Nest with "/": collection_name "AgentOPA/Formal" files the paper under an "AgentOPA" folder. MUTATES: also auto-saves the paper to the library. Idempotent (adding a paper already in the collection is a no-op). Requires SF_API_KEY.
Parameters3
arxiv_id
string
required
arXiv ID of the paper to add, e.g. '2407.15831'.
collection_name
string
optional
Name of the collection. Created if it doesn't exist. Use '/' to nest, e.g. 'AgentOPA/Formal'. Provide this OR collection_id.
collection_id
string
optional
UUID of an existing collection. Provide this OR collection_name.
Raw schema
{
"type": "object",
"properties": {
"arxiv_id": {
"type": "string",
"minLength": 1,
"description": "arXiv ID of the paper to add, e.g. '2407.15831'."
},
"collection_name": {
"description": "Name of the collection. Created if it doesn't exist. Use '/' to nest, e.g. 'AgentOPA/Formal'. Provide this OR collection_id.",
"type": "string",
"minLength": 1
},
"collection_id": {
"description": "UUID of an existing collection. Provide this OR collection_name.",
"type": "string"
}
},
"required": [
"arxiv_id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
remove_from_collection
Remove a paper from a collection, addressed by collection_id OR collection_name. MUTATES (the paper stays in your library; it's only removed from this collection). Idempotent. Requires SF_API_KEY.
Parameters3
arxiv_id
string
required
arXiv ID of the paper to remove from the collection.
collection_name
string
optional
Name of the collection. Provide this OR collection_id.
collection_id
string
optional
UUID of the collection. Provide this OR collection_name.
Raw schema
{
"type": "object",
"properties": {
"arxiv_id": {
"type": "string",
"minLength": 1,
"description": "arXiv ID of the paper to remove from the collection."
},
"collection_name": {
"description": "Name of the collection. Provide this OR collection_id.",
"type": "string",
"minLength": 1
},
"collection_id": {
"description": "UUID of the collection. Provide this OR collection_name.",
"type": "string"
}
},
"required": [
"arxiv_id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
create_watch
Create a standing watch β evaluated daily against newly-indexed papers, surfacing new matches via the email digest and via check_watches. MUTATES. Get-or-create by name (re-creating with an existing name returns it unchanged β never errors on duplicate). TWO forms: (1) the v2 STRUCTURED filter via `criteria` (collections/authors/categories/text/has_code/min_novelty/similar, AND-composed) β the composable, agent-tunable form, recommended; tune it with preview_watch first, and edit later with update_watch. Structured watches rank by 'rising' (forecasted breakout impact) by default, and tighten with min_impact_pct for an anti-noise watch that surfaces only the breakout papers in your niche. (2) a single legacy seed selector (q OR collection_name OR collection_id OR anchor_paper_id); if `criteria` is given it takes precedence. Requires SF_API_KEY.
Parameters11
name
string
required
Label for the watch, e.g. 'novel KV-cache work'.
novelty_min
number
optional
Only surface papers at/above this novelty score (0..1). The signal/noise knob β raise it for 'only tell me when it matters'. Default 0.5.
q
string
optional
Semantic/keyword topic seed. One seed selector only.
collection_name
string
optional
Watch the neighborhood of a collection by name (resolved by the backend). One seed selector only.
collection_id
string
optional
Watch the neighborhood of a collection by UUID. One seed selector only.
anchor_paper_id
string
optional
Watch papers similar to this arXiv ID. One seed selector only.
scope_to_citations_of
string
optional
Watch new papers citing this arXiv ID. One seed selector only.
author_id
string
optional
Watch an author's new work, by author ID. One seed selector only.
category
string
optional
Watch an arXiv category (e.g. 'cs.LG'), filtered by novelty_min. One seed selector only.
criteria
object
optional
v2 STRUCTURED filter (collections/authors/categories/text/has_code/min_novelty/similar). When provided, this defines the watch (kind='filter') and the single-selector seeds above are IGNORED. This is the composable, agent-tunable form β call preview_watch first to tune it.
recency_days
integer
optional
For a structured (criteria) watch: only consider papers from the last N days (default 7; the 'cites' relation uses 30).
Raw schema
{
"type": "object",
"properties": {
"name": {
"type": "string",
"minLength": 1,
"maxLength": 100,
"description": "Label for the watch, e.g. 'novel KV-cache work'."
},
"novelty_min": {
"default": 0.5,
"description": "Only surface papers at/above this novelty score (0..1). The signal/noise knob β raise it for 'only tell me when it matters'. Default 0.5.",
"type": "number",
"minimum": 0,
"maximum": 1
},
"q": {
"description": "Semantic/keyword topic seed. One seed selector only.",
"type": "string",
"minLength": 1
},
"collection_name": {
"description": "Watch the neighborhood of a collection by name (resolved by the backend). One seed selector only.",
"type": "string",
"minLength": 1
},
"collection_id": {
"description": "Watch the neighborhood of a collection by UUID. One seed selector only.",
"type": "string"
},
"anchor_paper_id": {
"description": "Watch papers similar to this arXiv ID. One seed selector only.",
"type": "string"
},
"scope_to_citations_of": {
"description": "Watch new papers citing this arXiv ID. One seed selector only.",
"type": "string"
},
"author_id": {
"description": "Watch an author's new work, by author ID. One seed selector only.",
"type": "string"
},
"category": {
"description": "Watch an arXiv category (e.g. 'cs.LG'), filtered by novelty_min. One seed selector only.",
"type": "string"
},
"criteria": {
"description": "v2 STRUCTURED filter (collections/authors/categories/text/has_code/min_novelty/similar). When provided, this defines the watch (kind='filter') and the single-selector seeds above are IGNORED. This is the composable, agent-tunable form β call preview_watch first to tune it.",
"type": "object",
"properties": {
"collections": {
"description": "Watch papers related to one or more of your collections.",
"type": "object",
"properties": {
"ids": {
"type": "array",
"items": {
"type": "string"
},
"description": "Collection UUIDs."
},
"relation": {
"default": "similar",
"description": "How a new paper relates to the collection(s): 'similar' = semantic neighborhood (broad); 'cites' = the new paper cites a collection member (uses a 30-day window); 'by_authors' = shares an author with the collection. NOTE: 'similar' here uses the default 0.70 cosine floor β to tighten it, ALSO pass a top-level `similar` predicate targeting the same collection (e.g. similar:{to:'collection:<that uuid>', min_score:0.9}); the collections group has no floor of its own.",
"type": "string",
"enum": [
"similar",
"cites",
"by_authors"
]
}
},
"required": [
"ids"
]
},
"authors": {
"description": "Match papers by these authors.",
"type": "object",
"properties": {
"names": {
"description": "Author display names.",
"type": "array",
"items": {
"type": "string"
}
},
"ids": {
"description": "Author UUIDs.",
"type": "array",
"items": {
"type": "string"
}
}
}
},
"categories": {
"description": "arXiv categories, e.g. ['cs.SE','cs.AI'].",
"type": "array",
"items": {
"type": "string"
}
},
"text": {
"description": "Keyword (full-text) or regex match on title/abstract.",
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Keyword/phrase (or a regex when mode='regex')."
},
"field": {
"type": "string",
"enum": [
"title",
"abstract",
"title_abstract"
]
},
"mode": {
"type": "string",
"enum": [
"fulltext",
"regex"
]
}
},
"required": [
"query"
]
},
"has_code": {
"description": "Only papers with released code.",
"type": "boolean"
},
"min_novelty": {
"description": "Novelty floor (0..1).",
"type": "number",
"minimum": 0,
"maximum": 1
},
"similar": {
"description": "Rank by semantic similarity to a target, with an optional cosine floor. Target the SAME collection used in `collections` (to:'collection:<uuid>') to put a floor on a collection-neighborhood watch.",
"type": "object",
"properties": {
"to": {
"type": "string",
"description": "Target: \"collection:<uuid>\" | \"paper:<arxivId>\" | \"text:<phrase>\"."
},
"min_score": {
"description": "Cosine floor (default 0.70).",
"type": "number"
}
},
"required": [
"to"
]
},
"rank": {
"description": "How to order matches. 'rising' (default) ranks by forecasted breakout impact first (the impact_pct momentum model), falling back to novelty when a paper is not impact-scored yet; 'novelty' ranks most-novel first; 'recent' ranks newest first; 'relevance' ranks by closeness to a similar target (needs a similar target, else behaves as rising). For a creator or stay-current watch, leave it as rising.",
"type": "string",
"enum": [
"rising",
"novelty",
"recent",
"relevance"
]
},
"min_impact_pct": {
"description": "Momentum floor: only surface papers in the top of forecasted citation impact within their field (e.g. 80 means roughly the top 20 percent). Only recently-scored papers have an impact percentile, so this also implies recent papers only β exactly right for a what-is-rising-now watch.",
"type": "integer",
"minimum": 0,
"maximum": 100
}
}
},
"recency_days": {
"description": "For a structured (criteria) watch: only consider papers from the last N days (default 7; the 'cites' relation uses 30).",
"type": "integer",
"minimum": 1,
"maximum": 60
}
},
"required": [
"name"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
list_watches
List the authenticated user's watches with name, a one-line definition summary, last_evaluated_at, and pending_hits (count of new matches since the last digest delivery). Read-only. Use before create_watch to see what's already tracked. Requires SF_API_KEY.
Pull new matching papers since the last digest delivery, in the same shape as search_papers results. Optionally scope to one watch by watch_name OR watch_id; omit both for all watches. Read-only and idempotent β does NOT advance any watermark (only digest delivery does), so it is safe to call repeatedly (no mark-on-read). This is the in-session 'anything new on my watches?' pull. Requires SF_API_KEY.
Parameters3
watch_name
string
optional
Scope to one watch by name. Provide this OR watch_id, or neither for all.
watch_id
string
optional
Scope to one watch by UUID. Provide this OR watch_name, or neither for all.
limit
integer
optional
Max hits to return (max 100).
Raw schema
{
"type": "object",
"properties": {
"watch_name": {
"description": "Scope to one watch by name. Provide this OR watch_id, or neither for all.",
"type": "string",
"minLength": 1
},
"watch_id": {
"description": "Scope to one watch by UUID. Provide this OR watch_name, or neither for all.",
"type": "string"
},
"limit": {
"default": 50,
"description": "Max hits to return (max 100).",
"type": "integer",
"minimum": 1,
"maximum": 100
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}
delete_watch
Delete a watch, addressed by watch_id OR name. MUTATES. Idempotent: deleting a non-existent watch is a no-op (no error). To change a watch in place (rename / novelty_min / retarget criteria) use update_watch instead of delete-and-recreate. Requires SF_API_KEY.
Parameters2
name
string
optional
Name of the watch to delete. Provide this OR watch_id.
watch_id
string
optional
UUID of the watch to delete. Provide this OR name.
Raw schema
{
"type": "object",
"properties": {
"name": {
"description": "Name of the watch to delete. Provide this OR watch_id.",
"type": "string",
"minLength": 1
},
"watch_id": {
"description": "UUID of the watch to delete. Provide this OR name.",
"type": "string"
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}
update_watch
Update an existing watch in place β rename, change novelty_min, or RETARGET its structured filter `criteria`. MUTATES. Address by watch_id OR name. Changing criteria replaces the definition and clears the watch's pending hits (so stale matches don't deliver); the next daily eval repopulates. Structured watches rank by 'rising' (forecasted breakout impact) by default, and tighten with min_impact_pct for an anti-noise watch that surfaces only the breakout papers in your niche. Tune the new criteria with preview_watch first. Requires SF_API_KEY.
Parameters6
name
string
optional
Find the watch by its current name. Provide this OR watch_id.
watch_id
string
optional
Find the watch by UUID. Provide this OR name.
new_name
string
optional
Rename the watch.
novelty_min
number
optional
New novelty floor (0..1).
criteria
object
optional
Replace the watch's filter (becomes kind='filter'). Clears pending hits.
recency_days
integer
optional
Window for the new criteria.
Raw schema
{
"type": "object",
"properties": {
"name": {
"description": "Find the watch by its current name. Provide this OR watch_id.",
"type": "string",
"minLength": 1
},
"watch_id": {
"description": "Find the watch by UUID. Provide this OR name.",
"type": "string"
},
"new_name": {
"description": "Rename the watch.",
"type": "string",
"minLength": 1,
"maxLength": 100
},
"novelty_min": {
"description": "New novelty floor (0..1).",
"type": "number",
"minimum": 0,
"maximum": 1
},
"criteria": {
"description": "Replace the watch's filter (becomes kind='filter'). Clears pending hits.",
"type": "object",
"properties": {
"collections": {
"description": "Watch papers related to one or more of your collections.",
"type": "object",
"properties": {
"ids": {
"type": "array",
"items": {
"type": "string"
},
"description": "Collection UUIDs."
},
"relation": {
"default": "similar",
"description": "How a new paper relates to the collection(s): 'similar' = semantic neighborhood (broad); 'cites' = the new paper cites a collection member (uses a 30-day window); 'by_authors' = shares an author with the collection. NOTE: 'similar' here uses the default 0.70 cosine floor β to tighten it, ALSO pass a top-level `similar` predicate targeting the same collection (e.g. similar:{to:'collection:<that uuid>', min_score:0.9}); the collections group has no floor of its own.",
"type": "string",
"enum": [
"similar",
"cites",
"by_authors"
]
}
},
"required": [
"ids"
]
},
"authors": {
"description": "Match papers by these authors.",
"type": "object",
"properties": {
"names": {
"description": "Author display names.",
"type": "array",
"items": {
"type": "string"
}
},
"ids": {
"description": "Author UUIDs.",
"type": "array",
"items": {
"type": "string"
}
}
}
},
"categories": {
"description": "arXiv categories, e.g. ['cs.SE','cs.AI'].",
"type": "array",
"items": {
"type": "string"
}
},
"text": {
"description": "Keyword (full-text) or regex match on title/abstract.",
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Keyword/phrase (or a regex when mode='regex')."
},
"field": {
"type": "string",
"enum": [
"title",
"abstract",
"title_abstract"
]
},
"mode": {
"type": "string",
"enum": [
"fulltext",
"regex"
]
}
},
"required": [
"query"
]
},
"has_code": {
"description": "Only papers with released code.",
"type": "boolean"
},
"min_novelty": {
"description": "Novelty floor (0..1).",
"type": "number",
"minimum": 0,
"maximum": 1
},
"similar": {
"description": "Rank by semantic similarity to a target, with an optional cosine floor. Target the SAME collection used in `collections` (to:'collection:<uuid>') to put a floor on a collection-neighborhood watch.",
"type": "object",
"properties": {
"to": {
"type": "string",
"description": "Target: \"collection:<uuid>\" | \"paper:<arxivId>\" | \"text:<phrase>\"."
},
"min_score": {
"description": "Cosine floor (default 0.70).",
"type": "number"
}
},
"required": [
"to"
]
},
"rank": {
"description": "How to order matches. 'rising' (default) ranks by forecasted breakout impact first (the impact_pct momentum model), falling back to novelty when a paper is not impact-scored yet; 'novelty' ranks most-novel first; 'recent' ranks newest first; 'relevance' ranks by closeness to a similar target (needs a similar target, else behaves as rising). For a creator or stay-current watch, leave it as rising.",
"type": "string",
"enum": [
"rising",
"novelty",
"recent",
"relevance"
]
},
"min_impact_pct": {
"description": "Momentum floor: only surface papers in the top of forecasted citation impact within their field (e.g. 80 means roughly the top 20 percent). Only recently-scored papers have an impact percentile, so this also implies recent papers only β exactly right for a what-is-rising-now watch.",
"type": "integer",
"minimum": 0,
"maximum": 100
}
}
},
"recency_days": {
"description": "Window for the new criteria.",
"type": "integer",
"minimum": 1,
"maximum": 60
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}
preview_watch
Dry-run a structured filter over recent papers WITHOUT creating a watch β the tuning loop. Returns {window_days, needs_similarity, match_count, sample} so you can iterate (add a category, raise min_novelty, switch the collection relation) before saving with create_watch. Structured watches rank by 'rising' (forecasted breakout impact) by default, and tighten with min_impact_pct for an anti-noise watch that surfaces only the breakout papers in your niche. NOTE: for a similarity filter, match_count is capped at 200 (the cosine fetch window) and so saturates at 200 on broad/hot topics β tune by the `sample` scores and narrow with categories/min_novelty (or a higher similar floor) rather than relying on match_count alone. Read-only. Requires SF_API_KEY.
Parameters2
criteria
object
required
The structured filter to test.
recency_days
integer
optional
Window in days (default 7; the 'cites' relation uses 30).
Raw schema
{
"type": "object",
"properties": {
"criteria": {
"type": "object",
"properties": {
"collections": {
"description": "Watch papers related to one or more of your collections.",
"type": "object",
"properties": {
"ids": {
"type": "array",
"items": {
"type": "string"
},
"description": "Collection UUIDs."
},
"relation": {
"default": "similar",
"description": "How a new paper relates to the collection(s): 'similar' = semantic neighborhood (broad); 'cites' = the new paper cites a collection member (uses a 30-day window); 'by_authors' = shares an author with the collection. NOTE: 'similar' here uses the default 0.70 cosine floor β to tighten it, ALSO pass a top-level `similar` predicate targeting the same collection (e.g. similar:{to:'collection:<that uuid>', min_score:0.9}); the collections group has no floor of its own.",
"type": "string",
"enum": [
"similar",
"cites",
"by_authors"
]
}
},
"required": [
"ids"
]
},
"authors": {
"description": "Match papers by these authors.",
"type": "object",
"properties": {
"names": {
"description": "Author display names.",
"type": "array",
"items": {
"type": "string"
}
},
"ids": {
"description": "Author UUIDs.",
"type": "array",
"items": {
"type": "string"
}
}
}
},
"categories": {
"description": "arXiv categories, e.g. ['cs.SE','cs.AI'].",
"type": "array",
"items": {
"type": "string"
}
},
"text": {
"description": "Keyword (full-text) or regex match on title/abstract.",
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Keyword/phrase (or a regex when mode='regex')."
},
"field": {
"type": "string",
"enum": [
"title",
"abstract",
"title_abstract"
]
},
"mode": {
"type": "string",
"enum": [
"fulltext",
"regex"
]
}
},
"required": [
"query"
]
},
"has_code": {
"description": "Only papers with released code.",
"type": "boolean"
},
"min_novelty": {
"description": "Novelty floor (0..1).",
"type": "number",
"minimum": 0,
"maximum": 1
},
"similar": {
"description": "Rank by semantic similarity to a target, with an optional cosine floor. Target the SAME collection used in `collections` (to:'collection:<uuid>') to put a floor on a collection-neighborhood watch.",
"type": "object",
"properties": {
"to": {
"type": "string",
"description": "Target: \"collection:<uuid>\" | \"paper:<arxivId>\" | \"text:<phrase>\"."
},
"min_score": {
"description": "Cosine floor (default 0.70).",
"type": "number"
}
},
"required": [
"to"
]
},
"rank": {
"description": "How to order matches. 'rising' (default) ranks by forecasted breakout impact first (the impact_pct momentum model), falling back to novelty when a paper is not impact-scored yet; 'novelty' ranks most-novel first; 'recent' ranks newest first; 'relevance' ranks by closeness to a similar target (needs a similar target, else behaves as rising). For a creator or stay-current watch, leave it as rising.",
"type": "string",
"enum": [
"rising",
"novelty",
"recent",
"relevance"
]
},
"min_impact_pct": {
"description": "Momentum floor: only surface papers in the top of forecasted citation impact within their field (e.g. 80 means roughly the top 20 percent). Only recently-scored papers have an impact percentile, so this also implies recent papers only β exactly right for a what-is-rising-now watch.",
"type": "integer",
"minimum": 0,
"maximum": 100
}
},
"description": "The structured filter to test."
},
"recency_days": {
"description": "Window in days (default 7; the 'cites' relation uses 30).",
"type": "integer",
"minimum": 1,
"maximum": 60
}
},
"required": [
"criteria"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
find_gaps
Find important work you HAVEN'T saved, for a collection or topic β a 'what am I missing?' analysis. Returns two buckets: foundational_gaps (canonical citation-graph anchors in the niche, not in your library) and frontier_gaps (recent high-novelty work in the niche, not yet saved). Provide exactly one seed: collection_name OR collection_id OR topic. The backend derives the niche, runs lineage + recent-novelty search, and subtracts your saved set. Read-only. Requires SF_API_KEY (it needs your library to subtract) and is a Pro feature β free accounts receive an upgrade prompt.
Parameters5
collection_name
string
optional
Analyze gaps for a collection by name (resolved by the backend). Provide exactly one seed.
collection_id
string
optional
Analyze gaps for a collection by UUID. Provide exactly one seed.
topic
string
optional
Analyze gaps for a free-text topic/area. Provide exactly one seed.
scope
string
optional
Which gaps to surface: 'foundational' (canonical anchors you're missing), 'frontier' (recent novel work you haven't saved), or 'both' (default).
limit
integer
optional
Max gaps per bucket (max 50). Default 10.
Raw schema
{
"type": "object",
"properties": {
"collection_name": {
"description": "Analyze gaps for a collection by name (resolved by the backend). Provide exactly one seed.",
"type": "string",
"minLength": 1
},
"collection_id": {
"description": "Analyze gaps for a collection by UUID. Provide exactly one seed.",
"type": "string"
},
"topic": {
"description": "Analyze gaps for a free-text topic/area. Provide exactly one seed.",
"type": "string",
"minLength": 1
},
"scope": {
"default": "both",
"description": "Which gaps to surface: 'foundational' (canonical anchors you're missing), 'frontier' (recent novel work you haven't saved), or 'both' (default).",
"type": "string",
"enum": [
"foundational",
"frontier",
"both"
]
},
"limit": {
"default": 10,
"description": "Max gaps per bucket (max 50). Default 10.",
"type": "integer",
"minimum": 1,
"maximum": 50
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}
ask_library
Answer a question using ONLY the papers you've saved β a synthesis over your library (or one collection) with inline [arXiv-ID] citations. The inverse of find_gaps (which finds important work you're MISSING): ask_library reasons over what you HAVE. Optionally scope to one collection (collection_name OR collection_id); omit both to use your whole library. Read-only. Requires SF_API_KEY (it reads your saved set). Free accounts get 1 question/month; Pro raises this to 200/day.
Parameters4
question
string
required
The natural-language question to answer from your saved papers.
collection_name
string
optional
Scope the answer to one collection by name (resolved by the backend). Omit to use your whole library.
collection_id
string
optional
Scope the answer to one collection by UUID. Omit to use your whole library.
limit
integer
optional
How many of your most-relevant saved papers to ground the answer on (max 20). Default 8.
Raw schema
{
"type": "object",
"properties": {
"question": {
"type": "string",
"minLength": 5,
"maxLength": 500,
"description": "The natural-language question to answer from your saved papers."
},
"collection_name": {
"description": "Scope the answer to one collection by name (resolved by the backend). Omit to use your whole library.",
"type": "string",
"minLength": 1
},
"collection_id": {
"description": "Scope the answer to one collection by UUID. Omit to use your whole library.",
"type": "string"
},
"limit": {
"default": 8,
"description": "How many of your most-relevant saved papers to ground the answer on (max 20). Default 8.",
"type": "integer",
"minimum": 1,
"maximum": 20
}
},
"required": [
"question"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
annotate_paper
Record YOUR verdict on a paper β why it matters for your work, when to use it, or why you ruled it out. One note per paper, upserted (writing again replaces it), so it is safe to call repeatedly. Requires SF_API_KEY. WHY IT MATTERS: this note is the only thing that survives between sessions. list_library returns note_text on every saved paper, so a verdict written now is what a future session reads INSTEAD of re-reading the paper and re-deriving the same conclusion. WRITE A JUDGMENT, NOT A SUMMARY β the paper already carries llm_summary and an abstract, so restating what the paper says adds nothing. Write what those cannot: how it bears on YOUR problem. Prefer a claim someone could later prove wrong ("needs a labeled trace log we don't have", "our baseline β beat this on the 7B setting") over an unfalsifiable verdict ("interesting", "not very relevant"), because a mechanism can be re-checked when circumstances change and a sentiment cannot. State the basis when it is thin: a verdict formed from the abstract alone deserves "(abstract only)", since fetch_fulltext defaults to ~800 characters of the results section rather than the whole paper. Pass action='get' to read the existing note before overwriting it β worth doing when a prior session may already have judged this paper. To correct a note, just write the corrected text (it replaces).
Parameters3
arxiv_id
string
required
arXiv ID of the paper to annotate, e.g. '2407.15831'.
note_text
string
optional
Your verdict (max 5000 chars). Required for the default upsert; ignored for action='get'.
action
string
optional
'upsert' (default) writes/replaces the note. 'get' returns the current note without changing it. There is deliberately no delete: a wrong note is corrected by overwriting it, which keeps this tool non-destructive.
Raw schema
{
"type": "object",
"properties": {
"arxiv_id": {
"type": "string",
"minLength": 1,
"description": "arXiv ID of the paper to annotate, e.g. '2407.15831'."
},
"note_text": {
"description": "Your verdict (max 5000 chars). Required for the default upsert; ignored for action='get'.",
"type": "string",
"minLength": 1,
"maxLength": 5000
},
"action": {
"default": "upsert",
"description": "'upsert' (default) writes/replaces the note. 'get' returns the current note without changing it. There is deliberately no delete: a wrong note is corrected by overwriting it, which keeps this tool non-destructive.",
"type": "string",
"enum": [
"upsert",
"get"
]
}
},
"required": [
"arxiv_id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}
check_drift
Answers 'for my problem, is the method I use superseded β and by what?' over a grounded, entity-resolved knowledge base of textual critique receipts + benchmark-dominance edges (no LLM call at query time). Call with a `family` (e.g. 'rag', 'peft', 'kvcache') and a `method` (e.g. 'SnapKV', 'LoRA') to get a verdict: how superseded it is, WHO critiques it (verbatim quotes + the citing paper), WHO beats it on benchmarks (winner, numbers, condition, source paper), and the not-yet-superseded alternatives in the same sub-problem. Omit `method` to get the whole-family map: most-superseded baselines, competition sub-problems, and the live frontier. Method names are matched case- and spacing-insensitively, with did-you-mean suggestions on a miss. Use this when choosing or reviewing a technique for a known problem area, or to check whether a baseline a paper relies on has been beaten. Does not require a Pro API key. Covers ~10 builder-problem families and growing; the `family` parameter lists them, or pass family='list' for the live set. Coverage caveat: evidence is drawn only from arXiv benchmark tables, so 'superseded' means a method was beaten in a published comparison (not that it is dead or unusable), production frameworks (LangChain, LlamaIndex, etc.) appear only as baselines and never as winners, and results are a literature signal rather than a deployment recommendation. GROUNDING β how far to trust an individual receipt: every claim passes a deterministic gate against the source paper's raw LaTeX (a critique must carry a verbatim quote shingle found in the source; a benchmark edge must have every one of its numbers present there), so a fabricated quote or table cell cannot enter the KB. What the gate does NOT verify is ATTRIBUTION: the quote is real but its subject may be class-level or a pronoun ('these methods', 'they') rather than the named method, so tying a receipt to one specific method is sometimes an inference. No end-to-end precision number has been measured on this endpoint β read the verbatim quote and its citing paper before repeating a verdict, and cite the source rather than asserting supersession as fact.
Parameters3
family
string
optional
Builder-problem family to query, e.g. 'rag' (retrieval-augmented generation), 'peft' (parameter-efficient fine-tuning), 'kvcache' (KV-cache compression). Omit it (or pass an unknown family like 'list') to get the live list of available families to pick from β start here if you don't know the family for a method.
method
string
optional
Method to check, e.g. 'SnapKV', 'H2O', 'StreamingLLM' (case/spacing-insensitive). Omit to get the whole-family map instead of a single-method verdict.
limit
integer
optional
Max items per list β receipts, dominance edges, frontier (3β50, default 12).
Raw schema
{
"type": "object",
"properties": {
"family": {
"description": "Builder-problem family to query, e.g. 'rag' (retrieval-augmented generation), 'peft' (parameter-efficient fine-tuning), 'kvcache' (KV-cache compression). Omit it (or pass an unknown family like 'list') to get the live list of available families to pick from β start here if you don't know the family for a method.",
"type": "string",
"minLength": 1,
"maxLength": 60
},
"method": {
"description": "Method to check, e.g. 'SnapKV', 'H2O', 'StreamingLLM' (case/spacing-insensitive). Omit to get the whole-family map instead of a single-method verdict.",
"type": "string",
"maxLength": 80
},
"limit": {
"default": 12,
"description": "Max items per list β receipts, dominance edges, frontier (3β50, default 12).",
"type": "integer",
"minimum": 3,
"maximum": 50
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}
Research paper search with ranking and citation tracking, for LLM engineering and academic research, without leaving Claude Code, Cursor, or any MCP client.
Most paper tools hand back a flat list. Scholar Feed ranks it: sort by relevance, by proven citation count, or by rising impact, then trace any paper's citation lineage forward and backward across 22M+ edges. 600k+ CS/AI/ML papers, updated daily, each with an LLM-generated summary and novelty score.
Scholar Feed indexes arXiv papers daily and ranks them on recency, citation velocity, institutional reputation, and code availability.
Quick Start
bash
npx scholar-feed-mcp@latest init
This interactive wizard will:
Optionally ask for an API key (or skip for anonymous access)
Detect your MCP client (Claude Code, Cursor, or Claude Desktop)
Write the config and verify the connection
No API key required. Anonymous access gives you 200 calls/month, enough for a typical research session. For a higher quota (500/month per account) plus your library β collections, saved papers and watches β get a free key at scholarfeed.org/settings.
Try asking: "Search for recent papers on test-time compute scaling"
What You Can Do
Technology scouting: "What novel research on retrieval-augmented generation was published this month?"
Literature review: "Find papers similar to 2401.04088 and export their BibTeX"
Trend monitoring: "What's trending in cs.CV this week? Summarize the top 3."
Author discovery: "Who are the top researchers working on efficient LLM inference?"
Field orientation: "Give me an orientation report on sparse mixture-of-experts architectures."
Installation
The fastest path is npx scholar-feed-mcp@latest init, which auto-detects your client and writes the config. To set it up by hand, every client launches the same stdio server (npx -y scholar-feed-mcp@latest); only the config-file location and the wrapper key differ.
Claude Desktop (one-click) installs without editing any config: download the .mcpb bundle from the latest release and open it (or drag it into Settings > Extensions). The installer shows one optional field for a Scholar Feed API key (sf_...): leave it blank for anonymous mode (200 calls/month), or paste a free key from scholarfeed.org/settings for 500/month.
Claude Code takes a one-line command:
bash
# Anonymous (200 calls/month)
claude mcp add scholar-feed -- npx -y scholar-feed-mcp@latest
# With an API key (500 calls/month per account)
claude mcp add scholar-feed -e SF_API_KEY=sf_your_key_here -- npx -y scholar-feed-mcp@latest
Every other client takes this standard JSON block:
To raise the quota to 500 calls/month, add "env": { "SF_API_KEY": "sf_your_key_here" } to the server entry. Get a free key at scholarfeed.org/settings.
Drop that block into the right config file:
Client
Config file
Notes
Cursor
.cursor/mcp.json (project) or ~/.cursor/mcp.json (global)
Program tab β Install β Edit mcp.json. Follows Cursor's notation.
JetBrains (PyCharm / IntelliJ)
AI Assistant β MCP β Add β As JSON
Requires AI Assistant 2025.1+.
A few clients need a different wrapper key or file format:
OpenAI Codex, VS Code (GitHub Copilot), Zed, Continue, and project-scoped configs
OpenAI Codex (~/.codex/config.toml, or $CODEX_HOME/config.toml if you set that) uses TOML, not JSON β the block above will not work. One file serves both the Codex CLI and the IDE extension.
Drop the env line to run keyless at 200 calls/month. On Windows, if Codex cannot launch the server, use command = "cmd" with args = ["/c", "npx", "-y", "scholar-feed-mcp@latest"].
VS Code: GitHub Copilot (.vscode/mcp.json) uses a servers key and an explicit type, and needs Copilot agent mode. You can also run MCP: Add Server from the Command Palette.
Get full paper details by arXiv ID. Also handles batch lookup and BibTeX export.
arxiv_ids, format, fields, verbose
get_citations
Citation graph (outgoing refs or incoming citations)
arxiv_id, direction, limit, fields
fetch_fulltext
Read a paper's text by section (abstract, introduction, related_work, method, results, conclusion, or all). Pass arxiv_ids to read up to 8 papers in one call; a paper that cannot be extracted comes back as a failed entry, not a failed call.
arxiv_id, arxiv_ids, sections
Authors
Tool
Description
Key Parameters
find_author
Find researchers by topic/name query, or retrieve a profile by ID.
q, id, field, limit
co_author_graph
Co-authorship neighborhood for an author
author_ids, window_years
Embeddings
Tool
Description
Key Parameters
embed_text
Get a 768-dim Gemini embedding for text (for HyDE and custom similarity). Pro-only, so anonymous/free callers get a 403 pro_required.
text, task_type
Research
Tool
Description
Key Parameters
get_field_orientation
Cheap retrieval orientation for a research area: top papers, subfields, open problems. No Pro quota.
topic, limit
get_foundational_lineage
Foundational work for a paper's niche via the citation graph (consensus-then-lift): niche_roots β field_level β discipline, with cited_by_in_niche evidence. Surfaces canonical anchors semantic search misses. No Pro quota.
anchor_paper_id, scope, generality_ceiling, limit
check_drift
"Is the method I use superseded β and by what?" Critique receipts + benchmark-dominance edges over ~10 LLM builder-problem families. No Pro quota.
family, method, limit
Library, Collections, Watches & Gap Analysis (require SF_API_KEY)
These MUTATE or read the authenticated user's account. The core read/search tools above work anonymously; these need a key.
Tool
Description
Key Parameters
save_paper
Bookmark a paper to your library (idempotent; feeds personalization).
arxiv_id
unsave_paper
Remove a paper from your library (idempotent).
arxiv_id
like_paper
"More like this" calibration signal for the For You feed (insert-only).
arxiv_id
list_library
List your saved papers, newest first (includes your notes).
limit, page
annotate_paper
Record your verdict on a paper β why it matters, when to use it, why you ruled it out. Upserted; returned by list_library, so it is what a later session reads instead of re-deriving.
arxiv_id, note_text, action
list_collections
List collections with paper counts.
(none)
create_collection
Create a named collection (get-or-create; no error on duplicate).
name
add_to_collection
Add a paper to a collection by name or id (also auto-saves).
arxiv_id, collection_name, collection_id
remove_from_collection
Remove a paper from a collection (stays saved).
arxiv_id, collection_name, collection_id
create_watch
Standing daily-evaluated saved search; get-or-create by name. Define it with a structured criteria filter (recommended) or a single seed selector.
"Answer from my saved set": a cited synthesis over your library or one collection, grounded only in papers you've saved (read-only). The inverse of find_gaps. Free 20/month, then Pro 200/day.
question, collection_name, collection_id, limit
Novelty Score
Every paper has an llm_novelty_score from 0.0 to 1.0:
Range
Meaning
Example
0.7+
Paradigm shift or broad SOTA
New architecture that changes the field
0.5-0.7
Novel method with strong results
New training technique with clear gains
0.3-0.5
Incremental improvement
Applying known method to new domain
<0.3
Survey, dataset, or minor extension
Literature review, benchmark release
Use novelty_min: 0.5 in search_papers to filter for genuinely novel work.
Rate Limits
Endpoint
Limit
search_papers
30/min
get_paper
30/min
get_citations
30/min
fetch_fulltext (single paper)
10/min
fetch_fulltext (batch, 2-8 papers)
6/min
find_author
20/min
co_author_graph
20/min
embed_text
30/min
get_field_orientation
20/min
get_foundational_lineage
20/min
find_gaps
20/min
ask_library
10/min
Responses include X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset headers.
Monthly volume quota (separate from the per-minute limits above, counted per account across all your keys): 200 calls/month anonymous (per IP), 500/month with a free key, 10,000/month on Pro. A smaller daily cap (100 / 200 / 2,000) sits underneath it as a burst guardrail so a runaway loop cannot spend a month in an hour; hitting it returns a 429 with scope: "burst" and leaves your monthly quota untouched. Read your remaining month from GET /v1/health (monthly_limit / usage_this_month) before a batch.
The AI synthesis tools have their own limits: ask_library is 20/month free, then 200/day on Pro; find_gaps is Pro-only (a 403 pro_required otherwise). embed_text needs an account of any tier β anonymous callers get a 403 account_required.
Example Response
search_papers with q: "attention mechanism" returns:
json
{"papers":[{"arxiv_id":"2401.04088","title":"Attention Is All You Need (But Not All You Get)","authors":["A. Researcher","B. Scientist"],"year":2024,"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","arxiv_url":"https://arxiv.org/abs/2401.04088","has_code":true,"github_url":"https://github.com/example/repo","citation_count":42,"rank_score":0.73,"llm_summary":"Proposes a sparse attention variant that reduces compute by 60% while matching dense attention accuracy on 5 benchmarks.","llm_novelty_score":0.55}],"total":1847,"page":1,"limit":20,"next_cursor":"eyJzIjogMC43MywgImlkIjogIjI0MDEuMDQwODgifQ=="}
Pass next_cursor back to get the next page (keyset pagination, which is more stable than page numbers for large result sets).
Environment Variables
Variable
Required
Default
Description
SF_API_KEY
No
(none)
Your Scholar Feed API key (starts with sf_). Without it, runs in anonymous mode (200 calls/month).
SF_API_BASE_URL
No
Production URL
Override API base URL
Development
bash
npm install
npm run build # Build to build/
npm run dev # Watch mode
npm run typecheck # Type check without emitting
npm test# Run tests
"Authentication failed: your SF_API_KEY is invalid"
The key may have been revoked. Generate a new one at scholarfeed.org/settings. Or remove the key to use anonymous mode.
"Rate limit exceeded" or "Anonymous daily limit exceeded"
Anonymous mode allows 200 calls/month. Get a free API key at scholarfeed.org/settings for 500 calls/month per account, plus your library.
Server shows as "failed" with no error β especially right after an update
The first launch (and the first launch after each new release) makes npx download the package. The published bin is a single self-contained file with no dependency tree to resolve, so this is fast β but on a slow link it can still outrun your client's start-up timeout, and the server then shows as "failed" with no detail. Fixes: (1) warm the cache by running it once in a terminal β npx -y scholar-feed-mcp@latest --version β then restart your client; (2) raise the MCP start-up timeout if your client supports it (Claude Code: MCP_TIMEOUT=60000). For the fastest, offline-capable launches, install once globally and point the config at it instead of npx:
bash
npm install -g scholar-feed-mcp
# then in your MCP config: "command": "scholar-feed-mcp", "args": []
Tool calls time out or fail silently
Ensure Node.js 18+ is installed (node --version). Older versions lack the native fetch API.
Stale npx cache
The config blocks above pin scholar-feed-mcp@latest, which re-resolves the newest version each launch. If you previously used an unpinned scholar-feed-mcp and are stuck on an old build: npx --yes scholar-feed-mcp@latest.
Windows: "command not found"
Use "command": "cmd" with "args": ["/c", "npx", "-y", "scholar-feed-mcp@latest"] in your MCP config.
About Scholar Feed
Scholar Feed is a research-discovery engine for computer science and AI/ML papers, founded in 2025. It indexes 600,000+ papers from arXiv β ranked by novelty, citation velocity, and relevance β with LLM-generated summaries, a citation graph, author profiles, and full-text extraction. It is available as a website, a public REST API, and a Model Context Protocol (MCP) server that AI agents can call directly. This package (scholar-feed-mcp) is the open-source MCP server.
Optional free API key from scholarfeed.org/settings. Anonymous callers get 200 requests/month; a free account raises that to 500/month and unlocks your library, and Pro is 10,000/month.