File size: 8,906 Bytes
2862b00
 
 
 
 
99c136c
 
596e05f
 
 
 
99c136c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2862b00
 
 
 
 
 
 
 
 
 
 
 
 
 
 
99c136c
596e05f
bfe6079
99c136c
2862b00
99c136c
 
 
2862b00
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c3b49d6
2862b00
 
bfe6079
2862b00
0947aaa
2862b00
 
 
bfe6079
 
 
 
 
0947aaa
c3b49d6
bfe6079
 
c3b49d6
bfe6079
c3b49d6
bfe6079
2862b00
 
 
bfe6079
f5dcf82
bfe6079
2862b00
0892c76
2862b00
 
 
 
 
 
 
 
 
bfe6079
f5dcf82
bfe6079
2862b00
0892c76
2862b00
 
 
 
 
 
 
c665dc8
c3b49d6
c665dc8
 
78b33e8
 
 
 
c665dc8
2862b00
aa00971
c3b49d6
aa00971
 
 
 
 
 
 
 
 
 
c3b49d6
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
"""
Model Context Protocol (MCP) for Paper2Agent

This codebase provides comprehensive bioinformatics analysis tools for transcriptomics data using the decoupler package. It supports bulk RNA-seq, single-cell RNA-seq, spatial transcriptomics, and pseudotime analysis workflows. The tools enable enrichment analysis, transcription factor activity scoring, pathway analysis, and cross-species gene mapping for functional genomics research.

DATASET DECISION TREE
═══════════════════════════════════════════════════════════════════════
STEP 0 (if data lives in GEO):
  β†’ decoupler_load_geo_series_matrix(url_or_path, condition_column=...)
  This downloads, parses, and saves an h5ad. Then proceed to STEP 1.

STEP 1 β€” Always start with: decoupler_inspect_data(adata_path)
  Check: is_integer, has_negatives, likely_raw_counts, n_obs, orientation_warning

STEP 2 β€” If orientation_warning is set:
  The matrix is transposed (genes as rows). Fix it before proceeding.

STEP 3 β€” Choose path based on inspect results:

  PATH A β€” Raw counts (likely_raw_counts=True)
    decoupler_load_and_filter_data
    β†’ decoupler_preprocess_data
    β†’ decoupler_differential_expression(method='deseq2')
    β†’ enrichment tools

  PATH B β€” Pre-normalized / log-transformed (data_type='log_expression' or 'log_ratio_microarray')
    Examples: GEO microarray, log-CPM, log-TPM, log-normalized scRNA-seq
    SKIP decoupler_preprocess_data entirely.
    Step 2a (microarray only β€” if features_look_like_probes=True):
      β†’ decoupler_collapse_probes_to_genes  # map probe IDs β†’ gene symbols before DE
      Skip this step for RNA-seq (features are already gene symbols).
    Step 2b:
      β†’ decoupler_differential_expression(method='ttest')   # no R required
      β†’ decoupler_differential_expression(method='limma')   # preferred if R available
    β†’ enrichment tools
    ⚠ Do NOT use method='deseq2' or decoupler_preprocess_data on pre-normalized data.

  PATH UNKNOWN β€” data_type='unknown' (heuristics inconclusive)
    Default to PATH B with method='ttest'. Ask the user to confirm data is not raw
    integer counts before proceeding. If it is, switch to PATH A with method='deseq2'.
═══════════════════════════════════════════════════════════════════════

This MCP Server contains tools extracted from the following tutorial files:
1. example
    - decoupler_load_toy_data: Load toy dataset and visualize gene expression patterns
    - decoupler_run_individual_methods: Run individual enrichment methods (GSEA and ULM)
    - decoupler_run_multiple_methods: Run all enrichment methods and compute consensus scores
    - decoupler_access_prior_knowledge: Query OmniPath resources for prior knowledge networks
2. licenses
    - decoupler_collect_transcriptional_regulators: Collect transcriptional regulator data with specified license
    - decoupler_access_resource_database: Access specific OmniPath resources with license restrictions
3. orthologs
    - decoupler_show_supported_organisms: Display organisms supported for ortholog translation
    - decoupler_load_knowledge_base: Load human-centric knowledge database resource
    - decoupler_translate_gene_symbols: Convert gene symbols from human to target organism
    - decoupler_access_organism_database: Directly load database with organism-specific gene symbols
4. rna
    - decoupler_inspect_data: Inspect data structure, detect orientation and data type, recommend analysis path
    - decoupler_load_geo_series_matrix: Download and parse a GEO series matrix URL or file into h5ad
    - decoupler_annotate_probes_with_gpl: Add gene symbols to var from a GPL platform SOFT file (required before probe collapse)
    - decoupler_collapse_probes_to_genes: Collapse microarray probe IDs to gene symbols before enrichment
    - decoupler_load_and_filter_data: Load bulk RNA-seq data and apply expression filtering
    - decoupler_preprocess_data: Normalize, scale and perform PCA analysis on bulk RNA-seq data (raw counts only)
    - decoupler_differential_expression: Differential expression via DESeq2, limma, or t-test (method parameter)
    - decoupler_tf_enrichment_collectri: TF enrichment using CollecTRI; auto-selects top TFs for network plot
    - decoupler_pathway_enrichment_progeny: Perform pathway enrichment analysis using PROGENy
    - decoupler_hallmark_enrichment: Perform hallmark gene set enrichment analysis
5. rna_pstime
    - decoupler_load_and_explore_data: Load and explore single-cell data with initial visualizations
    - decoupler_infer_pseudotime: Infer pseudotime trajectories using diffusion maps and DPT
    - decoupler_analyze_tf_activity: Analyze transcription factor activity along pseudotime
    - decoupler_analyze_pathway_activity: Analyze pathway activity using PROGENy along pseudotime
    - decoupler_analyze_hallmark_pathways: Analyze Hallmark pathway activity along pseudotime
6. rna_sc
    - decoupler_load_and_visualize_data: Load single-cell data and create initial UMAP visualization
    - decoupler_score_cell_types: Score cells using marker genes from PanglaoDB
    - decoupler_score_transcription_factors: Score cells using transcription factor networks from CollecTRI
    - decoupler_score_progeny_pathways: Score cells using PROGENy pathway gene sets
    - decoupler_score_hallmark_pathways: Score cells using Hallmark gene sets
7. rna_visium
    - decoupler_load_and_preprocess_visium: Load and preprocess spatial transcriptomics data with normalization
    - decoupler_apply_spatial_weighting: Apply spatial connectivity weighting to gene expression data
    - decoupler_score_transcription_factors: Score transcription factor activity using CollecTRI network
    - decoupler_score_progeny_pathways: Score pathway activity using PROGENy pathway database
    - decoupler_score_hallmark_pathways: Score pathway activity using Hallmark gene sets
"""

import os
import sys
import warnings

PROJECT_ROOT = os.path.dirname(os.path.abspath(__file__))
if PROJECT_ROOT not in sys.path:
    sys.path.insert(0, PROJECT_ROOT)

# ── R dependency installation ────────────────────────────────────────────────
# Run once at startup so limma is available when decoupler_differential_expression
# is called with method='limma'.  Failures are non-fatal: the rna.py tool
# handles a missing limma gracefully via fallback_to_ttest or RuntimeError.
try:
    from src.mcp_server.dependencies.r_dependencies import ensure_r_dependencies

    ensure_r_dependencies()
except FileNotFoundError as _e:
    warnings.warn(f"R install script not found, skipping R setup: {_e}", stacklevel=2)
except Exception as _e:
    warnings.warn(f"R dependency installation failed, limma may be unavailable: {_e}", stacklevel=2)

from fastmcp import FastMCP

# Import statements (alphabetical order)
from src.tools.bulk_dataset_tools import bulk_dataset_mcp
from src.tools.bulk_rnaseq import bulk_rnaseq_mcp
from src.tools.dataset_tools import dataset_mcp
from src.tools.example import example_mcp
from src.tools.integration_tools import integration_mcp
from src.tools.licenses import licenses_mcp
from src.tools.orthologs import orthologs_mcp
from src.tools.rna import rna_mcp
from src.tools.rna_pstime import rna_pstime_mcp
from src.tools.rna_sc import rna_sc_mcp
from src.tools.rna_visium import rna_visium_mcp

# Server definition and mounting
mcp = FastMCP(name="Paper2Agent")
mcp.mount(bulk_dataset_mcp)
mcp.mount(bulk_rnaseq_mcp)
mcp.mount(dataset_mcp)
mcp.mount(example_mcp)
mcp.mount(integration_mcp)
mcp.mount(licenses_mcp)
mcp.mount(orthologs_mcp)
mcp.mount(rna_mcp)
mcp.mount(rna_pstime_mcp)
mcp.mount(rna_sc_mcp)
mcp.mount(rna_visium_mcp)

import threading

from src.cache import preload_datasets

# Pre-populate the disk cache in a background thread (pure Python/numpy/scanpy,
# no rpy2 β€” safe to run concurrently with tool calls).
# R/limma installation is handled at module import above (already synchronous).
threading.Thread(target=preload_datasets, daemon=True, name="cache-preloader").start()

if __name__ == "__main__":
    import argparse

    parser = argparse.ArgumentParser(add_help=False)
    parser.add_argument("--transport", default=None)
    parser.add_argument("--port", type=int, default=8765)
    args, _ = parser.parse_known_args()

    if args.transport:
        # HTTP mode: persistent server, accepts multiple connections
        mcp.run(transport=args.transport, host="0.0.0.0", port=args.port)
    else:
        # stdio mode: default for MCP clients that spawn a subprocess per call
        mcp.run()