Abstract
Objective: Glioblastoma (GBM) and brain arteriovenous malformation (bAVM) are distinct cerebrovascular pathologies defined by aberrant vasculature. We investigated whether brain endothelial cells (ECs) in both conditions converge on a common transcriptional program carrying prognostic value in diffuse glioma.
Methods: We reanalyzed single-cell transcriptomes from 163,802 Raw FACS-purified human brain ECs across GBM, bAVM, lower-grade glioma (LGG), and non-pathological controls. Using patient-level pseudobulk differential expression, we derived a 100-gene shared module, assessed its stability (leave-onepatient-out, bootstrap), and adjusted for endothelial composition, age, and sex. Replication was performed in an independent single-cell EC cohort, bulk datasets, and 669 TCGA gliomas.
Results: Patient-level pseudobulk analysis identified a 100-gene disease-associated module (led by ESM1, ADAMTS5, PLVAP). Independently derived GBM-up and bAVM-up gene sets shared 605 genes (p = 7.3 × 10-293). Module induction persisted within matched endothelial subtypes across both diseases, showing a graded response: control < LGG < bAVM < GBM (p = 1.4 × 10-9). In patient-level models, the GBM effect survived full covariate adjustment (p = 0.0067), whereas the bAVM effect attenuated due to compositional shifts. The module replicated in independent single-cell (p = 4.8 × 10-19) and bulk glioma cohort, and higher expression independently predicted shorter overall survival in TCGA gliomas (adjusted HR = 1.38, p = 1.1 × 10-6).
Conclusion: GBM and bAVM share a common endothelial transcriptional state driven by vascular remodeling and inflammatory angiogenesis. In GBM, this state represents intrinsic endothelial reprogramming with significant prognostic value, whereas in bAVM it is largely linked to compositional changes.
Keywords: Brain endothelium; Glioblastoma; Single-cell RNA sequencing; Tumor vasculature.
1. Introduction
The cerebral vasculature is not a passive conduit. Brain endothelial cells (ECs) establish and maintain the blood-brain barrier (BBB), regulate immune trafficking, and instruct the surrounding neural and perivascular niche [1]. In disease, this specialized endothelium is remodeled, and the nature of that remodeling shapes the clinical phenotype [2]. Two conditions illustrate the extremes of cerebrovascular pathology. Glioblastoma (GBM), the most aggressive diffuse glioma [3,4], drives florid pathological angiogenesis, with microvascular proliferation as a defining histological hallmark, and its neovasculature is leaky, tortuous and functionally abnormal [5,6]. Brain arteriovenous malformation (bAVM) is a congenital, non-neoplastic tangle of direct artery-to-vein shunts lacking an intervening capillary bed, and is a leading cause of hemorrhagic stroke in young adults [7,8]. One is a cancer; the other is a developmental vascular anomaly. Mechanistically they are treated as unrelated.
A dysfunctional, structurally abnormal endothelium with disrupted arteriovenous identity, breakdown of the normal capillary/BBB program, and prominent extracellular-matrix (ECM) remodeling [9]. This raises a question that single‑cell transcriptomics is now positioned to answer directly, namely whether the endothelia of GBM and bAVM converge on a shared molecular program, despite arising from entirely different pathogenic origins. The existence of such a program would suggest that the brain endothelium possesses a limited, stereotyped repertoire of disease‑associated states, analogous to reactive astrocytosis [10,11] or the recently described disease‑associated microglia [12], that different pathologies independently engage.
Single-cell transcriptomics has resolved the vascular cell types of the brain, first in mouse [13] and subsequently in human tissue, where atlases of the normal, Alzheimer-affected and malformed cerebrovasculature defined arteriovenous zonation and disease-associated shifts [14–16]. Two recent resources make the direct GBM–bAVM comparison feasible for the first time. Wälchli et al. (Nature 2024) generated a single-cell atlas of the human brain vasculature across development, adulthood and disease (GEO GSE256493) [17], including FACS-purified CD31⁺/CD45⁻ endothelial cells from GBM, bAVM, lower-grade glioma (LGG) and non-pathological control brain. These datasets contain the same purified populations, processed identically, across all entities. Independently, Xie et al. (Neuron 2024) profiled the cerebrovasculature of brain malignancies and matched normal tissue (GSE242044) [18], providing an orthogonal cohort for replication. Together they allow a directly comparable, cross-entity dissection of the diseased brain endothelium.
We test whether GBM neovasculature and brain AVM share a disease-associated endothelial gene program. We (i) define the module by patient-level pseudobulk differential expression of GBM versus non-pathological control endothelium, with per-gene stability quantified by leave-one-patient-out and bootstrap resampling, an analysis unit adopted because cell-level testing inflates false discoveries in multi-patient designs [19-21]; (ii) test whether it survives patient-balanced downsampling and adjustment for endothelial subtype composition, age and sex; (iii) score it across bAVM, LGG and control and separate compositional remodeling from within-subtype reprogramming using subtype-matched contrasts; (iv) derive the gene program shared between independently computed GBM and bAVM up-regulated sets; (v) scrutinize the mesenchymal component for cell-identity ambiguity and test specificity against expression-matched nulls and negativecontrol gene sets; (vi) replicate in two independent glioma cohorts; (vii) characterize the biology by pathway enrichment; and (viii) ask whether the program carries prognostic value in 669 TCGA gliomas. We are explicit throughout about which claims the data support and which they do not.
2. Methods
2.1 Data Sources
Primary discovery cohort (Wälchli et al., GSE256493): We obtained the authors’ pre-processed, FACS-sorted (CD31⁺/CD45⁻) endothelial-cell Seurat objects from the GEO supplementary files for four entities: glioblastoma, brain arteriovenous malformation, lower-grade glioma, and non-pathological temporal-lobe control tissue. Each object carried the authors’ own cell-level annotations, including endothelial subtype (ECclusters), pathology label, patient identifier, sex and age. Raw endothelial-cell counts were 49,999 (GBM, 8 patients), 20,305 (bAVM, 5 patients), 17,373 (LGG, 6 patients) and 76,125 (control, 9 patients). These raw counts reproduced the published figures exactly, confirming object identity. Cohort composition, quality-control outcomes, and the downsampling scheme are summarized in Table 1.
Table 1. Endothelial-cell cohort (Wälchli GSE256493, FACS-purified CD31⁺/CD45⁻).
Entity Patients Raw ECs After QC ECs analyzed Downsampled to Glioblastoma (GBM)
| Entity | Patients | Raw ECs | After QC | ECs analyzed | Downsampled to |
|---|---|---|---|---|---|
| Glioblastoma (GBM) | 8 | 49,999 | 49,593 | 48,157 | 6,000 |
| Brain AVM (bAVM) | 5 | 20,305 | 20,208 | 19,667 | 6,000 |
| Lower-grade glioma (LGG) | 6 | 17,373 | 17,314 | 16,970 | 6,000 |
| Control (temporal lobe) | 9 | 76,125 | 75,650 | 74,919 | 6,000 |
| Combined (integrated) | 28 | — | — | — | 24,000 |
Independent replication cohort (Xie et al., GSE242044): Raw 10x Genomics count matrices for 13 samples (6 glioma, 7 histologically normal cerebrovasculature) were downloaded from GEO (GSE242044_RAW.tar) and processed de novo (see 2.6).
External prognostic cohort (TCGA): RNA-seq V2 expression and overlaid clinical/survival data for the TCGA PanCancer Atlas lower-grade glioma (lgg_tcga_pan_can_atlas_2018) and glioblastoma (gbm_tcga_pan_can_atlas_2018) cohorts were retrieved programmatically from cBioPortal.
2.2 Quality control and cohort construction
All single-cell analysis used R with Seurat v5.5.1. Because each sorted object expands to several gigabytes in memory, each entity was processed in an isolated R subprocess: raw counts and slim metadata were extracted directly, qualitycontrol (QC) metrics were computed at the matrix level, and a fresh object was rebuilt to avoid carrying the authors’ precomputed graphs. Cells were retained with 200 ≤ detected genes ≤ 8,000, mitochondrial fraction ≤ 20%, and non-zero total counts; cells annotated as a technical “Mitochondrial” cluster were removed. To prevent any single large entity (GBM, control) from dominating the integrated embedding, each entity was randomly downsampled to a fixed cap of 6,000 ECs (set.seed (1234)), yielding a balanced combined cohort of 24,000 ECs × 26,666 genes spanning 13 endothelial subtypes. All patient-level (pseudobulk) analyses used the full, non-downsampled cohort.
2.3 Integration, clustering and subtype annotation
The combined object was log-normalized, and 2,000 variable features were identified, scaled, and reduced by PCA (30 components). Batch effects across entities were corrected with Seurat canonical-correlation analysis integration (CCAIntegration). A shared-nearest-neighbor graph (dims 1–30) was clustered at resolution 0.5 (13 clusters) and embedded by UMAP on the integrated reduction. We used the authors’ own endothelial subtype labels and, for interpretability, consolidated them onto the canonical arteriovenous zonation axis (arterial, capillary, venous, EndoMT, proliferating, stem-to-EC). Subtype identity was confirmed against 17 canonical markers: GJA5/BMX/SEMA3G (arterial), MFSD2A/SLC7A5/CD320 (capillary/BBB), NR2F2/ACKR1/SELP (venous), ACTA2/TAGLN/COL1A1 (EndoMT), MKI67/TOP2A (proliferating) and PECAM1/CLDN5/VWF (pan-endothelial).
2.4 Patient-level pseudobulk as the primary discovery analysis
Provenance of all differential-expression inference: Every differential-expression result reported here is computed on raw RNA counts (DefaultAssay = “RNA”, Seurat v5 layers joined with JoinLayers), never on integrated or corrected values. The CCA-integrated reduction (2.3) is used only for visualization (UMAP) and for graph-based clustering; no statistical inference is performed on it. This addresses the concern that batch-corrected expression is inappropriate for hypothesis testing.
Pseudobulk construction: For each patient, raw counts were summed across that patient’s endothelial cells to a single library (Matrix: rowSums over the RNA counts layer), producing a genes × patient matrix (28 patients: GBM 8, bAVM 5, LGG 6, control 9). A second, finer aggregation summed counts per patient × endothelial subtype for the subtype-resolved analyses (2.6).
Module definition: GBM-versus-control differential expression was performed with edgeR: DGEList → filterByExpr (group = entity) → TMM normalization → estimateDisp → glmQLFit → glmQLFTest. Genes were called up-regulated at FDR < 0.05, log₂FC > 1 and logCPM > 0, yielding 732 genes. The primary module is the top 100 of these ranked by log₂FC × (−log₁₀ FDR). For the pre-revision cell-level module (Wilcoxon FindMarkers, logfc.threshold = 0.10, min.pct = 0.10, only.pos = TRUE, on log-normalized RNA counts), see 3.2 and Supplementary Table S3. This module is retained only for comparison and for the TCGA survival analysis (2.12), which predates the redefinition.
2.5 Module stability and downsampling sensitivity.
Leave-one-patient-out (LOPO): The full pseudobulk pipeline was re-run 17 times, each time omitting one of the 17 GBM or control patients, and the top-100 module recomputed. For each gene we report the LOPO entry frequency, the fraction of runs in which it entered the top 100.
Bootstrap: Patients were resampled with replacement within group 200 times; the module was recomputed each time and per-gene bootstrap entry frequency recorded.
Core module: Genes with LOPO frequency = 1.0 and bootstrap frequency ≥ 0.8 constitute the stability-filtered core (n = 31). Per-gene LOPO and bootstrap frequencies for the 100 module genes are in Supplementary Table S4. Patient-balanced downsampling: To exclude the possibility that patients contributing many cells (25-2,632 per patient; Supplementary Table S1) drive the result, cells were downsampled to 150 per patient within patients under five random seeds, the full pseudobulk pipeline re-run for each, and cross-seed concordance of log₂FC assessed by Spearman correlation.
2.6 Differential composition and subtype-resolved analysis Composition:
Per-patient subtype proportions were centered-log-ratio transformed and compared between entities by PERMANOVA on Aitchison distances (999 permutations, vegan::adonis2), globally and per disease-vs-control contrast. Individual subtype proportions were compared by permutation test (999 label permutations) with Benjamini–Hochberg correction across subtype × contrast.
Subtype-matched differential expression: Using the patient × subtype pseudobulk matrix, edgeR contrasts were run within each of the three major subtypes (arterial, capillary, venous) for each disease versus control, retaining only patients contributing ≥ 10 cells to that subtype. This separates a change in cell-type mixture from a change in transcriptional state within a fixed cell type.
Composition adjustment: Patient-level module scores were modeled as score ~ entity (unadjusted), + age, + subtype proportions (arterial, capillary, venous, EndoMT), and fully adjusted (+ age + sex + proportions) by ordinary least squares.
2.7 Module scoring and cross-entity comparison
Module activity per cell was quantified with Seurat AddModuleScore (seed = 1234) on log-normalized RNA counts. Because cell-level tests over thousands of cells are anticonservative, all reported statistical comparisons use patient-level mean scores: per-patient means were compared to control by two-sided Wilcoxon rank-sum and Welch t-tests. Withinsubtype scores were computed by averaging cell-level scores within each patient × subtype cell.
2.8 Mesenchymal-signal scrutiny and specificity controls
Contamination assessment: For the EndoMT-annotated cluster we computed per-cell module scores for endothelial (PECAM1, CLDN5, VWF, CDH5, FLT1), mesenchymal (COL1A1, COL1A2, FN1, ACTA2, TAGLN, POSTN), mural (PDGFRB, RGS5, NOTCH3, KCNJ8), fibroblast (DCN, LUM, COL3A1) and immune (PTPRC) programs, and recorded per-marker detection rates. Cells were flagged as failing endothelial identity if they lacked detectable PECAM1 and CLDN5. The entire GBM-vs-control pseudobulk analysis was then repeated after excluding (i) the EndoMT cluster and (ii) all flagged cells, and log₂FC concordance with the full analysis assessed.
Expression-matched null: One thousand null modules of 100 genes each were drawn matched to the observed module on mean logCPM decile, and the observed median log₂FC compared to the null distribution (empirical p, z-score).
Unrelated gene sets: Housekeeping (ACTB, GAPDH, TUBB, RPL13A, n = 15) and canonical EC-identity genes (PECAM1, CLDN5, VWF, CDH5, n = 11) were scored as negative controls: a genuinely disease-associated program should move while these do not.
Graded-response test: Patient-level module scores were correlated with an ordinal severity rank (control < LGG < bAVM < GBM) by Spearman correlation.
2.9 Independent external validation (GSE16011)
The redefined module and its core were scored in an independent bulk glioma microarray cohort (GSE16011; 276 gliomas with WHO grade, 8 non-tumor reference samples) as the mean per-gene z-score across samples (93/100 primary and 29/31 core genes mappable). Association with WHO grade was tested by Spearman correlation and GBM-versus-lowergrade separation by one-sided Mann–Whitney U.
2.10 Independent replication (Xie cohort)
Each Xie 10x sample was processed with Seurat (min.cells = 3, min.features = 200; identical QC). After normalization, HVG selection, PCA and clustering (resolution 0.6), endothelial clusters were identified by mean expression of canonical EC markers (PECAM1/CLDN5/VWF/FLT1/CDH5/KDR) together with per-cell PECAM1/CLDN5 detection, yielding 37,286 tumor ECs. Because per-sample module scoring centers each sample against its own background and thereby cancels a shared signal, replication was assessed at the pseudobulk level: EC counts were aggregated per sample and tested (tumor vs normal, edgeR). Given the modest sample size, we evaluated coordinated directionality of the module, namely the fraction of module genes up-regulated in tumor endothelium (binomial test) and the shift of module-gene log₂FC relative to background (Wilcoxon), rather than per-gene significance.
2.11 Pathway enrichment
The 605-gene shared program was tested for over-representation against MSigDB Hallmark, Reactome (C2:CP) and Gene Ontology Biological Process (C5) collections by the hypergeometric test, using the set of tested genes as the background universe (the appropriate background for a differential-expression-derived gene list). Gene sets were restricted to 10–500 members and required ≥ 3 overlapping genes; p-values were Benjamini–Hochberg adjusted.
2.12 Survival analysis
For each TCGA glioma patient, a module score was computed as the mean, across the 99 mappable module genes, of the per-gene z-score of log₂ (RSEM + 1) expression. Overall survival was modeled with Cox proportional-hazards regression on the z-standardized score (hazard ratio per standard deviation), in the pooled pan-glioma cohort (LGG + GBM), within each grade separately, and adjusted for grade and age. Kaplan–Meier curves used score tertiles with a log-rank test. Analyses used the survival and survminer packages.
2.13 Reproducibility
All figures were computed directly from the analysis objects; every quantitative statement in this manuscript derives from the accompanying supplementary tables. Random seeds were fixed for downsampling and module scoring. Analysis code and intermediate objects (integrated Seurat object, DE tables, module definitions, scoring tables) are provided as supplementary artifacts.
3. Results
3.1 A balanced, subtype-resolved atlas of the diseased brain endothelium
After QC, we assembled a balanced cohort of 24,000 FACS-purified brain endothelial cells (6,000 each from GBM, bAVM, LGG and control; 26,666 genes) and integrated them across entities by canonical-correlation analysis. The integrated UMAP resolved all four entities in a shared embedding while preserving endothelial heterogeneity (Figure 1A). Using the authors’ subtype annotations projected onto the arteriovenous axis, the atlas recovered the expected zonation (arterial, capillary, venous) together with an endothelial-to-mesenchymal transition (EndoMT) compartment, proliferating ECs and a small stem-to-EC population (Figure 1B). Subtype identity was confirmed by all 17 canonical markers, each expressed in its expected compartment (Figure 1C): arterial GJA5/BMX/SEMA3G, capillary/BBB MFSD2A/SLC7A5/CD320, venous NR2F2/ACKR1/SELP, mesenchymal ACTA2/TAGLN/COL1A1, proliferative MKI67/TOP2A, and pan-endothelial PECAM1/CLDN5/VWF. Entity composition differed as expected: control endothelium was dominated by capillary and arteriolar cells, whereas both GBM and bAVM showed expansion of angiogenic-capillary, mesenchymal and venous compartments (Supplementary Table S1).

Figure 1: A balanced, subtype-resolved single-cell atlas of the diseased human brain endothelium. (A) UMAP of 24,000 integrated FACS-purified brain ECs colored by entity (control, LGG, GBM, bAVM). (B) The same embedding colored by endothelial subtype consolidated onto the arteriovenous zonation axis. (C) Dot plot of 17 canonical endothelial markers across subtypes (dot size = fraction of cells expressing; color = scaled mean expression), confirming subtype identities.
For continuity with the pre-revision analysis, the earlier cell-level Wilcoxon module (2,479 up- and 1,120 down-regulated genes) shares 62 of 100 members with the patient-level module; the leading genes are common to both while the module tail differs, as expected when the unit of analysis changes. Reassuringly, all 100 patient-level module genes are also significantly up-regulated in the full-cohort pseudobulk analysis (median log₂FC 3.89), and genome-wide log₂FC ranks agree closely between the analysis cohort and the full cohort (Spearman ρ = 0.969 over 11,041 genes). All results below use the patient-level module.

Figure 2: Definition and cross-entity activity of the disease-associated endothelial module. (A) Volcano plot of GBM-vs-control endothelial differential expression (single-cell Wilcoxon, pre-revision analysis retained for comparison); module genes are highlighted and the top 15 labeled. The primary module definition is patient-level pseudobulk. (B) Module score across entities; violins show the cell-level distribution and overlaid points the per-patient means.
Downsampling to 150 cells per patient within patient under five random seeds reproduced the effect estimates closely (pairwise log₂FC Spearman ρ = 0.941–0.946; Figure 5b; Supplementary Table S5). Top-100 membership was less stable across seeds (31–42 genes overlapping the primary module), which is the expected consequence of ranking a discrete list under a 24-fold reduction in depth; the underlying effect estimates, not the list boundary, are what the module represents. We therefore report both the 100-gene module and the 31-gene core rather than treating any single 100-gene list as canonical.
3.4 Compositional remodeling and within-subtype reprogramming are separable
Endothelial subtype composition differs markedly between entities (global Aitchison PERMANOVA F = 3.32, R² = 0.293, p = 0.0048; Supplementary Table S6). GBM versus control is the strongest contrast (R² = 0.342, FDR = 0.0039), driven by arterial depletion (median proportion 0.13 vs 0.34, FDR = 0.006) and stem-to-EC expansion (FDR = 0.014); bAVM shows a 3.8-fold venous expansion (0.39 vs 0.10, FDR = 0.006) (Figure 5c; Supplementary Table S6a). A shared “program” could in principle be nothing more than this shift in cell-type mixture.
In subtype-matched contrasts built from the patient × subtype pseudobulk matrix, module induction persists within a fixed cell type (Supplementary Table S7). For GBM versus control in capillary endothelium alone, all 52 testable module genes are up-regulated (52 significantly; binomial p = 2.2×10⁻¹⁶; median log₂FC 4.42), including ADAMTS5 (log₂FC 11.36), ESM1 (10.97) and PLVAP (7.98). The same holds in arterial (20/20 up, 19 significant) and venous (38/38 up, 37 significant) compartments, and for bAVM versus control in all three subtypes (capillary 50/50 up, 48 significant, binomial p = 8.9×10⁻¹⁶; arterial 40/40, 35 significant; venous 35/36, 33 significant). Within-subtype module scores are elevated for both GBM and bAVM across capillary, arterial and venous compartments (all FDR < 0.004; Supplementary Table S8).
In patient-level models, the GBM effect is attenuated but not abolished by adjustment: β = 0.275 (p = 1.1×10⁻⁷) unadjusted, 0.217 (p = 0.00095) adjusting for subtype composition, and 0.202 (p = 0.0067) fully adjusted for composition, age and sex (Figure 5d; Supplementary Table S9). The bAVM effect, significant unadjusted (β = 0.147, p = 0.0022), falls to non-significance after composition adjustment (β = 0.083, p = 0.39) and remains so fully adjusted (β = 0.074, p = 0.46). We therefore distinguish the two diseases explicitly: in GBM the shared program reflects genuine within-subtype
transcriptional reprogramming that survives full adjustment; in bAVM the patient-level signal is statistically inseparable from compositional remodeling, even though the subtype-matched gene-level contrasts remain strongly concordant. With five bAVM patients, this analysis is underpowered to separate the two, and we do not claim it does. Taken together, these results demonstrate the size of the disease-associated module, its stability under resampling, and the separability of compositional and transcriptional effects are summarized in Table 2.
Table 2: Key quantitative results (revised analysis) Analysis Statistic Value Cells per patient range (median) 25–2,632 (744) GBM vs control, patient-lev pseudobulk (n=8 vs 9)
| Analysis | Statistic | Value |
|---|---|---|
| Cells per patient range (median) | range (median) | 25–2,632 (744) |
| GBM vs control, patient-level pseudobulk (n=8 vs 9) | significantly up-regulated genes | 732 |
| Module stability (core) | LOPO frequency = 1.0 / bootstrap ≥ 0.8 | 63 / 31 genes |
| Patient-balanced downsampling (5 seeds, 150 cells/patient) | pairwise log₂FC Spearman ρ | 0.941–0.946 |
| Endothelial composition | global Aitchison PERMANOVA R²; p | 0.293; 0.0048 |
| GBM vs control, capillary-only contrast | module genes up (binomial p) | 52 / 52 (2.2×10⁻¹⁶) |
| Module score, GBM vs control | β unadjusted → fully adjusted; p | 0.275 → 0.202; 0.0067 |
| Module score, bAVM vs control | β unadjusted → fully adjusted; p | 0.147 → 0.074; 0.46 |
| GBM vs bAVM up-set overlap | shared genes; fold over chance; hypergeometric p | 605; 4.5×; 7.3×10⁻²⁹³ |
| Genome-wide fold-change concordance | Spearman r | 0.52 |
| EndoMT-cluster identity | DCN / RGS5 / PDGFRB vs PECAM1 detection | 84.4% / 76.2% / 66.9% vs 54.9% |
| EndoMT + non-EC exclusion sensitivity | log₂FC correlation; module genes still significant | r = 0.991; 94 / 94 |
| Expression-matched null modules (n=1,000) | observed vs null median log₂FC; z | 3.82 vs 0.014; 69.9 |
| Negative controls | housekeeping / EC-identity median log₂FC | 0.30 / −0.22 |
| Severity gradient (control<LGG<bAVM<GBM) | Spearman ρ; p | 0.873; 1.4×10⁻⁹ |
| Shared-program enrichment | top Hallmark (adj p) | EMT, 8.5×10⁻²⁸ |
| Xie replication (redefined module) | genes up in tumor; binomial p | 81 / 86; 4.8×10⁻¹⁹ |
| GSE16011 bulk validation (n=276) | grade Spearman ρ; p | 0.49; 3.5×10⁻¹⁸ |
| TCGA survival (pre-revision module, n=669) | Cox HR per SD (95% CI); p | 1.34 (1.19–1.51); 1.5×10⁻⁶ |
| TCGA survival (grade+age-adjusted) | Cox HR per SD; p | 1.38; 1.1×10⁻⁶ |
TCGA survival (grade+age-adjusted) Cox HR per SD; p 1.38; 1.1×10⁻⁶
3.5 The program is shared by brain AVM endothelium
The central question was whether bAVM endothelium (a non-neoplastic, congenital lesion) engages the same program. As shown in figure 2A, 2B at the patient level, the GBM-derived module was significantly elevated in bAVM relative to control (mean per-patient score −0.042 vs −0.189; Wilcoxon p = 0.0010, Welch t-test p = 0.0019), with GBM highest (+0.085; Wilcoxon p = 8.2×10⁻⁵) and LGG intermediate (−0.106; Wilcoxon p = 0.018, Welch t-test p = 0.13, significant by rank test only) (Supplementary Table S8). As set out in 3.4, the bAVM patient-level elevation does not survive adjustment for compositional differences, so we frame it as a shared state observed in bAVM endothelium rather than an independently established, composition-independent induction.
To characterize the shared program without circularity, we computed GBM-up and bAVM-up gene sets independently (each vs control, patient-level pseudobulk edgeR, FDR < 0.05 and log₂FC > 1): 1,243 genes for GBM and 1,651 for bAVM. Within the 13,232-gene universe tested in both analyses (1,128 GBM-up and 1,582 bAVM-up genes fall inside it), their overlap was 605 genes, 4.5-fold above the 155 expected by chance (hypergeometric p = 7.3×10⁻²⁹³; Jaccard =
0.287), and genome-wide per-gene log₂FC values were strongly correlated between the two diseases (Spearman r = 0.52; Figure 3A). Of the 100 module genes, 99 were testable in the bAVM contrast and 83 were also independently upregulated in bAVM (median bAVM log₂FC 2.81); for the 31-gene core, 30 of 30 testable genes (median 3.33). The shared program was led by ESM1, CYTL1, ADAMTS5, COL15A1, POSTN, CCL20, FAP, STC1, SERPINA5, PLVAP, THBS1 and COL3A1, a mixture of angiogenic, matrix-remodeling and inflammatory mediators.
Correction to the submitted version: The overlap p-value reported at submission (≈5×10⁻²⁴⁹) was computed with unrestricted marginal counts against a restricted universe. Recomputed with marginals restricted to the shared universe, the correct value is 7.3×10⁻²⁹³. The shared-gene count (605), the correlation (r = 0.52) and every downstream conclusion are unchanged.

Figure 3: A gene program shared between glioblastoma and brain-AVM endothelium. (A) Per-gene log₂ fold-change (disease vs control) in GBM (x) versus bAVM (y), computed independently by patient-level pseudobulk. The 605 shared up-regulated genes are highlighted; annotation gives the genome-wide Spearman correlation and overlap statistics. (B) Over-representation analysis of the 605-gene shared program against MSigDB Hallmark, Reactome and GO:BP (hypergeometric test vs the tested-gene background; bar length = -log₁₀ adjusted p; labels = overlap/set size).
Critically, the module does not depend on these cells. Repeating the entire GBM-versus-control pseudobulk analysis after excluding the EndoMT cluster and all cells failing endothelial identity (lacking both PECAM1 and CLDN5; 16.9% of EndoMT-cluster cells) leaves the results essentially unchanged: log₂FC correlation with the full analysis r = 0.991, and 94 of 94 testable module genes remain significantly up-regulated (Supplementary Table S11). The shared program is a property of unambiguous endothelium.
3.7 Specificity: the program is not a generic stress or abundance artifact
Three controls argue that the module is a specific disease-associated signal rather than an artifact of expression level or generic tissue stress (Figure 6; Supplementary Table S12). First, against 1,000 expression-matched null modules drawn to match the observed logCPM distribution, the observed median log₂FC (3.82) exceeds the null median (0.014) by z = 69.9 (empirical p = 0.001, the resolution limit of 1,000 permutations); the core module gives z = 46.6. Second, negative-control gene sets do not move: housekeeping genes show a median log₂FC of 0.30 and canonical EC-identity genes −0.22, against
3.82 for the module, the endothelium is not globally shifted, and it has not simply lost identity. Third, the response is graded with disease severity rather than binary: patient-level module scores increase monotonically control < LGG < bAVM < GBM (Spearman ρ = 0.873, p = 1.4×10⁻⁹), and GBM exceeds LGG significantly (Wilcoxon p = 0.008).
3.8 The shared program is a mesenchymal, matrix-remodeling, inflammatory-angiogenic state
Over-representation analysis of the 605-gene shared program returned to highly coherent biology across three independent databases (Figure 3B). Among MSigDB Hallmark sets (24 significant), the strongest was epithelial–mesenchymal transition (52/172 genes, adjusted p = 8.5×10⁻²⁸), with the endothelial equivalent being EndoMT, followed by TNFα signaling via NF-κB, angiogenesis (11/30, p = 2.7×10⁻⁷), hypoxia, inflammatory response and coagulation. Reactome (78 significant) was dominated by extracellular-matrix organization (57/223, p = 2.1×10⁻²⁵), collagen formation, ECM proteoglycans and vascular-wall interactions; GO: BP (575 significant) reinforced external-encapsulating-structure/ECM organization, blood-vessel morphogenesis, collagen-fibril organization and wound response. Together these define the shared endothelium as an activated, partially mesenchymal state that has lost quiescent capillary/BBB identity in favor of matrix deposition, angiogenesis and inflammatory signaling.
3.9 Independent replication in two external cohorts
Single-cell replication (Xie cohort): We tested the redefined module in the independent Xie et al. glioma vascular atlas (GSE242044), processed de novo from raw counts (37,286 tumor ECs across 6 tumor and 7 normal samples). Because per-sample scoring centers each sample against its own background and cancels a shared signal (tumor vs normal persample means were indistinguishable, Wilcoxon p = 1.0), we assessed replication at the pseudobulk level. While no individual gene reached genome-wide significance at this sample size, the module replicated as a coordinated unit: 81 of 86 testable module genes were up-regulated in tumor endothelium (binomial p = 4.8×10⁻¹⁹; median log₂FC 1.97 vs −0.007 for background genes; Wilcoxon p = 6.2×10⁻⁴⁵) (Figure 7a, 7b). The 31-gene core behaves likewise (24 of 26 testable genes up, binomial p = 5.3×10⁻⁶; median log₂FC 2.34, Wilcoxon p = 6.8×10⁻¹⁴). We report this as module-level coordinated directionality, not per-gene replication; the cohort is too small for the latter.
Bulk replication (GSE16011): In an independent bulk glioma cohort of 276 tumors with WHO grade (93/100 module genes mappable), the module score increased monotonically with grade (Spearman ρ = 0.49, p = 3.5×10⁻¹⁸), separated GBM from lower-grade glioma (Mann–Whitney p = 8.7×10⁻¹⁵, n = 159 vs 117), and was markedly lower in the eight nontumor reference samples (mean −0.73 vs 0.02) (Figure 7b; Supplementary Table S13). The 31-gene core reproduces this (ρ = 0.39, p = 3.4×10⁻¹¹). This is an endothelium-derived signature measured in bulk tissue, so it necessarily reflects both endothelial state and vascular content; it provides consistent evidence, but not an independent endothelial measurement.
Note on bAVM: Neither replication cohort contains an arteriovenous malformation. The bAVM arm of this study is therefore based on a single cohort of five patients and has not been independently replicated, and the title and conclusions have been revised accordingly.
3.10 The module is prognostic in glioma, independent of grade
In 669 TCGA gliomas with matched RNA-seq and survival data (510 LGG, 159 GBM; cBioPortal PanCancer Atlas), a higher per-patient module score predicted shorter overall survival: pan-glioma Cox HR = 1.34 per standard deviation (95% CI 1.19–1.51; p = 1.5×10⁻⁶), with clear separation of survival by module tertile (Figure 4A, 4B; log-rank p = 4.2×10⁻⁶). Because glioma survival is dominated by grade, we confirmed the association within grade: within LGG the module was strongly prognostic (HR = 1.57 per SD; p = 2.8×10⁻⁹) and it remained significant within GBM (HR = 1.22; p = 0.037). In a model adjusting for both grade and age, the module retained independent prognostic value (HR = 1.38 per SD; p = 1.1×10⁻⁶) (Figure 4C). The endothelial disease program thus stratifies glioma outcome beyond established clinical variables.
Important caveat: This survival analysis was computed with the pre-revision, cell-level module definition. The TCGA expression matrix was not retained after the original analysis and it could not be re-accessed in the revision environment, so we could not recompute survival with the redefined patient-level module or its 31-gene core. Given that the two definitions share their leading genes and that the redefined module tracks WHO grade in independent bulk data (3.9), we expect the prognostic association to hold, but we have not demonstrated it and do not claim it. This is stated in the abstract and Limitations.

Figure 4: Independent replication and prognostic validation. (A) In the Xie cohort, distribution of tumor-vs-normal endothelial log₂ fold-changes for the original module genes versus all background genes. The equivalent panel for the redefined module is Figure 7a. (B) Kaplan–Meier overall survival for 669 TCGA gliomas stratified by module-score tertile. (C) Forest plot of Cox hazard ratios per SD of module score across models (pan-glioma, grade+age-adjusted, within-LGG, within-GBM).

Figure 5: Module robustness to patient resampling, cell-count imbalance and compositional confounding. (a) Per-gene bootstrap entry frequency versus leave-one-patient-out entry frequency for all 732 significantly up-regulated genes; the 31 core genes are highlighted. (b) Per-patient endothelial subtype proportions by entity, with permutation-test results for differentially abundant subtypes. (c) Median module log₂FC within endothelial subtypes across disease entities. (d) Patient-level module-score effect estimates β (95% CI) for each entity across unadjusted, age-adjusted, composition-adjusted and fully adjusted models.

Figure 6: Mesenchymal-signal scrutiny and specificity controls. (a) Marker detection rates in the EndoMT-annotated cluster, showing high mural and fibroblast marker positivity alongside endothelial markers. (b) Concordance of gene-level log₂FC before and after excluding the EndoMT cluster and all cells failing endothelial identity. (c) Median log₂FC of the module, core, housekeeping and canonical EC-identity gene sets. (d) Patient-level module scores across the severity gradient control < LGG < bAVM < GBM.

Figure 7: External validation of the redefined module. (a) In the Xie cohort, tumor-vs-normal endothelial log₂ fold-changes for the redefined module versus background. (b) Module score by WHO grade in the independent GSE16011 bulk glioma cohort (n = 276). (c) Median full-cohort pseudobulk log₂FC for the original cell-level module, the redefined patient-level module and the stability-filtered core, computed on one common matrix.
4. Discussion
We find that two mechanistically unrelated cerebrovascular diseases, glioblastoma, a malignant neoplasm, and brain arteriovenous malformation, a congenital vascular anomaly, converge on a shared disease-associated endothelial transcriptional state. A 100-gene module defined at the patient level from glioblastoma endothelium is elevated in brainAVM endothelium, and gene sets computed independently for each disease overlap by 605 genes, 4.5-fold above chance, with strongly correlated genome-wide fold-changes. The state is mesenchymal-like, matrix-remodeling and inflammatoryangiogenic; it is stable to patient resampling, survives adjustment for endothelial composition in GBM, and replicates in two independent glioma cohorts [18,22].
A stereotyped disease-endothelial state: The convergence of cancer and a developmental malformation on one endothelial program supports the idea that the brain endothelium, like astrocytes and microglia [10,12], has a bounded repertoire of reactive states rather than an unlimited, disease-specific response. The dominant axis of that state is a mesenchymal-like activation coupled to extracellular-matrix deposition. EMT-related and ECM/collagen organization sets were the strongest enrichments in two independent databases. We are deliberately careful with terminology here. Endothelial-tomesenchymal transition is well documented in cardiovascular disease, cerebral cavernous malformation and glioblastoma vasculature [23–27], but full transition implies a lineage conversion that transcriptome data cannot establish, and our scrutiny of the mesenchymal compartment (3.6) found that a substantial fraction of those cells carry pericyte and fibroblast markers at rates exceeding their endothelial markers, consistent with transition, but equally with mural-cell capture or ambient-RNA contamination. We note that neither ambient-RNA correction nor doublet detection was applied to the deposited objects we re-analyzed, and dedicated tools for both [28,29] would be the appropriate next step. Excluding those cells entirely leaves the module intact, so the mesenchymal-like signal is a property of unambiguous endothelium; what remains untested is whether individual cells are transitioning. Lineage tracing or dual-marker protein imaging would settle it. Layered on top are hypoxia, TNFα/NF-κB inflammation and coagulation, features consistent with the leaky, thrombosis-prone, inflamed vasculature seen histologically in both conditions [6,30].
Composition versus state: To separate two explanations that single-cell disease comparisons routinely conflate, a distinction that requires compositional as well as expression analysis [31]. Endothelial subtype composition does differ between entities, GBM loses arterial endothelium and gains stem-to-EC cells, bAVM expands its venous compartment nearly four-fold. But in GBM the module remains induced inside each matched subtype, with every testable module gene upregulated in capillary-versus-capillary contrasts, and the patient-level effect survives adjustment for composition, age and sex. In bAVM the gene-level within-subtype concordance is equally strong, yet the patient-level score effect does not survive composition adjustment with five patients. We now report these two diseases asymmetrically rather than asserting a uniform mechanism. GBM shows within-subtype reprogramming; bAVM shows a shared state entangled with compositional remodeling.
Recurrent driver genes: The leading shared genes are individually informative. ESM1 (endocan) is a VEGF-induced tipcell gene [23] and PLVAP a marker of barrier breakdown in experimental models [32]. Together they index the fenestrated, BBB-deficient angiogenic endothelium characteristic of GBM and are candidate imaging or therapeutic handles. CD276 (B7-H3), a module member and an actively pursued immunotherapy target in CNS tumors [33], shows disease-associated vascular expression that may matter for delivery. FAP, POSTN, THBS1, COL15A1 and ADAMTS5 implicate a matrix-remodeling, fibroblast-like endothelial phenotype, and both FAP and POSTN have independently been proposed as glioma stromal targets [34,35]. CCL20 and the NF-κB program point to endothelial participation in immune recruitment. That these emerge from both diseases nominates them as targets common to malignant and non-malignant cerebrovascular pathology.
Prognostic significance: The module’s independence from grade in TCGA glioma is notable: rather than simply tracking the LGG-to-GBM axis, the endothelial state carries prognostic information within grade, most strikingly within LGG, where outcomes are otherwise heterogeneous (HR = 1.57 per SD, p = 2.8×10⁻⁹). This positions the tumor endothelium, not only the tumor cells, as a candidate determinant of glioma outcome, complementing the molecular classification that now dominates glioma prognostication [36,37] and consistent with spatial evidence that vascular niches organize glioblastoma’s tumor host interface [38]. Two qualifications matter. The analysis uses the pre-revision module definition and could not be recomputed for the redefined module (3.10). Further it applies an endothelium-derived signature to bulk tumor RNA, where the endothelial fraction is diluted. The association survives that dilution, but bulk expression of endothelial genes reflects vascular content as well as endothelial state, so the finding is an association rather than evidence that endothelial reprogramming cause worse outcome.
Methodological rigor: Several choices guard against common single-cell pitfalls. We used identically FACS-purified
endothelium across all entities, removing the confound of variable EC capture. All differential expression is computed on raw RNA counts, never on integrated values; the CCA embedding is used only for visualization and clustering. Discovery is patient-level pseudobulk on the full cohort, so cells are not treated as independent replicates, and module membership is accompanied by explicit per-gene resampling frequencies rather than presented as a fixed list. Compositional and withinsubtype analyses separate mixture change from state change. The shared program was built from independently derived GBM and bAVM sets, so the overlap is not circular, and pathway enrichment used the tested-gene background. Where an analysis could not be performed, such as a developmental reference or other injury endothelia for recomputed survival, we explicitly state this limitation rather than substituting a weaker proxy. The analysis rests on established, openly available tooling and resources- Seurat v5 for integration and scoring [39], edgeR for pseudobulk differential expression [40], the MSigDB Hallmark collection for enrichment [41], cBioPortal for the TCGA survival cohort [42] and the survival package for Cox modeling [43].
4.1 Limitations
Several caveats should be stated plainly. We did not test whether this program is developmentally “reactivated,” as fetal brain endothelial data were inaccessible during revision, and this claim has been withdrawn from the title and text rather than argued from indirect evidence. Control endothelium derives from non-pathological temporal lobe, whereas GBM, LGG and bAVM lesions arise across varied brain regions, and the groups differ in age (median 25 versus 60 years for GBM); moreover, lesion location, surgical indication, ischemic interval and tissue-handling metadata are not deposited with the source dataset, so we cannot exclude that part of the disease–control difference reflects anatomical origin rather than pathology, with the within-subtype analyses serving as the only sensitivity check available. Cohorts are small and unbalanced (GBM n = 8, bAVM n = 5), per-patient cell yields vary widely (25–2,632 cells), and the bAVM compositionadjustment analysis is underpowered, so its null result should not be read as evidence of absence. Xie-cohort replication demonstrates module-level coordinated directionality rather than per-gene significance; GSE16011 is bulk tissue, confounding endothelial transcriptional state with vascular content; and neither external cohort contains bAVM, leaving the bAVM findings entirely unreplicated. Specificity is only partially established because we could not test endothelium from stroke, brain metastasis or non-CNS tumors, and we therefore cannot exclude that the program represents a general injured-endothelium response rather than one specific to GBM and bAVM. The TCGA survival analysis was performed with the pre-revision module definition and could not be recomputed for the redefined module or its 31-gene core. Finally, anti-angiogenic therapy has not improved overall survival in newly diagnosed glioblastoma, so this program should not be read as endorsing VEGF blockade; any therapeutic value lies in the specific surface genes nominated, which require prospective protein-level and functional validation, and mesenchymal-like activation together with every candidate target are inferred from RNA only, leaving cell-identity ambiguity in the mesenchymal compartment resolvable only by lineage tracing or dual-marker protein imaging.
5. Conclusion
Glioblastoma neovasculature and brain arteriovenous malformation share a reproducible, disease-associated endothelial gene program defined by mesenchymal-like activation, extracellular-matrix remodeling and inflammatory angiogenesis. A patient-level module derived from glioblastoma endothelium is elevated in brain-AVM endothelium; the two diseases’ independently derived up-regulated programs overlap 4.5-fold above chance; and the module is stable to patient resampling, persists within matched endothelial subtypes, and replicates in two independent glioma cohorts. In glioblastoma this reflects within-subtype transcriptional reprogramming that survives adjustment for cellular composition; in brain AVM five patients, one cohort, no independent replication the state is present but statistically inseparable from compositional remodeling. Whether the program recapitulates developmental angiogenesis, and whether it distinguishes these diseases from other injured endothelia, remain open. Subject to those caveats, the shared genes led by ESM1, PLVAP, ADAMTS5, FAP and POSTN nominate candidate endothelial markers spanning malignant and non-malignant cerebrovascular disease.
6. Declarations
Ethical approval
Ethical approval was not required for this study because it involved secondary analysis of publicly available, de-identified datasets. No new human participants were recruited, no biological samples were collected, and no identifiable patient information was accessed.
Consent
All participants provided voluntary written informed consent prior to participation, and no personally identifiable data were collected.
Sources of funding
This research received no external funding.
Author’s contribution
Conceptualization: Yahya Mahmood, Nadia Ishtiaq, Hadia Ishtiaq and Umair Ahmed; Methodology: Yahya Mahmood, Nadia Ishtiaq, Hadia Ishtiaq and Umair Ahmed; Data collection: Yahya Mahmood, Nadia Ishtiaq; Formal analysis and data curation: Yahya Mahmood, Nadia Ishtiaq, Hadia Ishtiaq and Umair Ahmed; Writing original draft preparation: Yahya Mahmood, Nadia Ishtiaq and Hadia Ishtiaq; Writing review and editing: Umair Ahmed; Supervision: Umair Ahmed.
All authors have read and approved the final version of the manuscript.
Conflicts of interest disclosure
The authors declare no conflicts of interest.
Data availability statement
All primary data analyzed are publicly available: the discovery brain-vascular atlas at GEO GSE256493 (Wälchli et al.), the replication cohort at GEO GSE242044 (Xie et al.), and the TCGA glioma PanCancer Atlas cohorts via cBioPortal (lgg_tcga_pan_can_atlas_2018, gbm_tcga_pan_can_atlas_2018). Analysis code and derived tables are available from the corresponding author upon reasonable request. Analyses used R with Seurat v5.5.1, edgeR/limma, and the survival/survminer packages.
Acknowledgement
The authors thank the original investigators of the publicly available datasets used in this study, specifically the Wälchli et al. brain vascular atlas (GEO GSE256493), the Xie et al. glioma vascular atlas (GEO GSE242044), and The Cancer Genome Atlas (TCGA) for generating and openly sharing these data resources.
AI Statement
The authors declare that generative AI (ChatGPT plus) was used solely for language polishing. All scientific content, interpretation, and final approval of the manuscript were performed by the authors.
7. Supplementary materials
Supplementary Materials
- Table S1 —
S1_cells_per_patient.csv: endothelial cells contributed by each patient, overall and by subtype. - Table S2 —
S2_cohort_characteristics.csv: per-patient age, sex, region, cell yield and module score, by entity. - Table S3 —
S3_module_pseudobulk_primary.csv: 100-gene patient-level pseudobulk module with edgeR statistics. - Table S4 —
S4_module_stability.csv: per-gene leave-one-patient-out and bootstrap entry frequencies; core-module flag. - Table S5 —
S5_patient_balanced_downsampling.csv: effect estimates and concordance across five downsampling seeds. - Table S6 —
S6a_composition_tests.csv,S6b_patient_subtype_proportions.csv,S6c_permanova.csv: differential composition results. - Table S7 —
S7_subtype_specific_DE.csv,S7b_module_within_subtype.csv: subtype-specific differential expression and module analysis within subtypes. - Table S8 —
S8_within_subtype_scores.csv: within-subtype module scores and patient-level cross-entity comparisons. - Table S9 —
S9_composition_adjusted_models.txt: unadjusted, age-adjusted, composition-adjusted, and fully adjusted model summaries. - Table S10 —
S10_endomt_marker_profile.csv: marker detection and program scores in the mesenchymal compartment. - Table S11 —
S11_endomt_excluded_sensitivity.csv: differential expression after excluding mesenchymal and non-EC cells. - Table S12 —
S12_negative_controls.csv: expression-matched null and negative-control gene-set results. - Table S13 —
S13_gse16011_external_validation.csv: per-sample module scores and grade in the bulk validation cohort.
References
- Sweeney MD, Zhao Z, Montagne A, Nelson AR, Zlokovic BV. Blood-brain barrier: from physiology to disease and back. Physiological Reviews. 2018 Oct 3.
- Daneman R, Prat A. The blood-brain barrier. Cold Spring Harbor Perspectives in Biology. 2015 Jan;7(1):a020412.
- Louis DN, Perry A, Wesseling P, Brat DJ, Cree IA, Figarella-Branger D, Hawkins C, Ng HK, Pfister SM, Reifenberger G, Soffietti R. The 2021 WHO classification of tumors of the central nervous system: a summary. Neuro-Oncology. 2021 Aug 1;23(8):1231-51.
- Price M, Ballard C, Benedetti J, Neff C, Cioffi G, Waite KA, Kruchko C, Barnholtz-Sloan JS, Ostrom QT. CBTRUS statistical report: primary brain and other central nervous system tumors diagnosed in the United States in 2017–2021. Neuro-Oncology. 2024 Oct 6;26(Suppl 6):vi1.
- Jain RK, Di Tomaso E, Duda DG, Loeffler JS, Sorensen AG, Batchelor TT. Angiogenesis in brain tumours. Nature Reviews Neuroscience. 2007 Aug;8(8):610-22.
- Hardee ME, Zagzag D. Mechanisms of glioma-associated neovascularization. The American Journal of Pathology. 2012 Oct 1;181(4):1126-41.
- Lawton MT, Rutledge WC, Kim H, Stapf C, Whitehead KJ, Li DY, Krings T, TerBrugge K, Kondziolka D, Morgan MK, Moon K. Brain arteriovenous malformations. Nature Reviews Disease Primers. 2015 May 28;1(1):15008.
- Solomon RA, Connolly Jr ES. Arteriovenous malformations of the brain. New England Journal of Medicine. 2017 May 11;376(19):1859-66.
- Carmeliet P, Jain RK. Molecular mechanisms and clinical applications of angiogenesis. Nature. 2011 May 19;473(7347):298-307.
- Liddelow SA, Guttenplan KA, Clarke LE, Bennett FC, Bohlen CJ, Schirmer L, Bennett ML, Münch AE, Chung WS, Peterson TC, Wilton DK. Neurotoxic reactive astrocytes are induced by activated microglia. Nature. 2017 Jan 26;541(7638):481-7.
- Escartin C, Galea E, Lakatos A, O’Callaghan JP, Petzold GC, Serrano-Pozo A, Steinhäuser C, Volterra A, Carmignoto G, Agarwal A, Allen NJ. Reactive astrocyte nomenclature, definitions, and future directions. Nature Neuroscience. 2021 Mar;24(3):312-25.
- Keren-Shaul H, Spinrad A, Weiner A, Matcovitch-Natan O, Dvir-Szternfeld R, Ulland TK, David E, Baruch K, Lara-Astaiso D, Toth B, Itzkovitz S. A unique microglia type associated with restricting development of Alzheimer’s disease. Cell. 2017 Jun 15;169(7):1276-90.
- Vanlandewijck M, He L, Mäe MA, Andrae J, Ando K, Del Gaudio F, Nahar K, Lebouvier T, Laviña B, Gouveia L, Sun Y. A molecular atlas of cell types and zonation in the brain vasculature. Nature. 2018 Feb 22;554(7693):475-80.
- Yang AC, Vest RT, Kern F, Lee DP, Agam M, Maat CA, Losada PM, Chen MB, Schaum N, Khoury N, Toland A. A human brain vascular atlas reveals diverse mediators of Alzheimer’s risk. Nature. 2022 Mar 31;603(7903):885-92.
- Garcia FJ, Sun N, Lee H, Godlewski B, Mathys H, Galani K, Zhou B, Jiang X, Ng AP, Mantero J, Tsai LH. Single-cell dissection of the human brain vasculature. Nature. 2022 Mar 31;603(7903):893-9.
- Winkler EA, Kim CN, Ross JM, Garcia JH, Gil E, Oh I, Chen LQ, Wu D, Catapano JS, Raygor K, Narsinh K. A single-cell atlas of the normal and malformed human brain vasculature. Science. 2022 Mar 4;375(6584):eabi7377.
- Wälchli T, Ghobrial M, Schwab M, Takada S, Zhong H, Suntharalingham S, Vetiska S, Gonzalez DR, Wu R, Rehrauer H, Dinesh A. Single-cell atlas of the human brain vasculature across development, adulthood and disease. Nature. 2024 Aug 15;632(8025):603-13.
- Xie Y, Yang F, He L, Huang H, Chao M, Cao H, Hu Y, Fan Z, Zhai Y, Zhao W, Liu X. Single-cell dissection of the human blood-brain barrier and glioma blood-tumor barrier. Neuron. 2024 Sep 25;112(18):3089-105.
- Squair JW, Gautier M, Kathe C, Anderson MA, James ND, Hutson TH, Hudelle R, Qaiser T, Matson KJ, Barraud Q, Levine AJ. Confronting false discoveries in single-cell differential expression. Nature Communications. 2021 Sep 28;12(1):5692.
- Zimmerman KD, Espeland MA, Langefeld CD. A practical solution to pseudoreplication bias in single-cell studies. Nature Communications. 2021 Feb 2;12(1):738.
- Crowell HL, Soneson C, Germain PL, Calini D, Collin L, Raposo C, Malhotra D, Robinson MD. Muscat detects subpopulation-specific state transitions from multi-sample multi-condition single-cell transcriptomics data. Nature Communications. 2020 Nov 30;11(1):6077.
- Gravendeel LA, Kouwenhoven MC, Gevaert O, de Rooi JJ, Stubbs AP, Duijm JE, Daemen A, Bleeker FE, Bralten LB, Kloosterhof NK, De Moor B. Intrinsic gene expression profiles of gliomas are a better predictor of survival than histology. Cancer Research. 2009 Dec 1;69(23):9065-72.
- Rocha SF, Schiller M, Jing D, Li H, Butz S, Vestweber D, Biljes D, Drexler HC, Nieminen-Kelhä M, Vajkoczy P, Adams S. Esm1 modulates endothelial tip cell behavior and vascular permeability by enhancing VEGF bioavailability. Circulation Research. 2014 Aug 29;115(6):581-90.
- Kovacic JC, Dimmeler S, Harvey RP, Finkel T, Aikawa E, Krenning G, Baker AH. Endothelial to mesenchymal transition in cardiovascular disease: JACC state-of-the-art review. Journal of the American College of Cardiology. 2019 Jan 22;73(2):190-209.
- Piera-Velazquez S, Jimenez SA. Endothelial to mesenchymal transition: role in physiology and in the pathogenesis of human diseases. Physiological Reviews. 2019 Apr 1;99(2):1281-324.
- Maddaluno L, Rudini N, Cuttano R, Bravi L, Giampietro C, Corada M, Ferrarini L, Orsenigo F, Papa E, Boulday G, Tournier-Lasserve E. EndMT contributes to the onset and progression of cerebral cavernous malformations. Nature. 2013 Jun 27;498(7455):492-6.
- Huang M, Liu T, Ma P, Mitteer RA, Zhang Z, Kim HJ, Yeo E, Zhang D, Cai P, Li C, Zhang L. c-Met–mediated endothelial plasticity drives aberrant vascularization and chemoresistance in glioblastoma. The Journal of Clinical Investigation. 2016 May 2;126(5):1801-14.
- Young MD, Behjati S. SoupX removes ambient RNA contamination from droplet-based single-cell RNA sequencing data. GigaScience. 2020 Dec;9(12):giaa151.
- McGinnis CS, Murrow LM, Gartner ZJ. DoubletFinder: doublet detection in single-cell RNA sequencing data using artificial nearest neighbors. Cell Systems. 2019 Apr 24;8(4):329-37.
- Potente M, Carmeliet P. The link between angiogenesis and endothelial metabolism. Annual Review of Physiology. 2017 Feb 10;79:43-66.
- Buettner M, Ostner J, Mueller CL, Theis FJ, Schubert B. scCODA is a Bayesian model for compositional single-cell data analysis. Nature Communications. 2021 Nov 25;12(1):6876.
- Shue EH, Carson-Walter EB, Liu Y, Winans BN, Ali ZS, Chen J, Walter KA. Plasmalemmal vesicle associated protein-1 (PV-1) is a marker of blood-brain barrier disruption in rodent models. BMC Neuroscience. 2008 Feb 26;9(1):29.
- Majzner RG, Theruvath JL, Nellan A, Heitzeneder S, Cui Y, Mount CW, Rietberg SP, Linde MH, Xu P, Rota C, Sotillo E. CAR T cells targeting B7-H3, a pan-cancer antigen, demonstrate potent preclinical activity against pediatric solid tumors and brain tumors. Clinical Cancer Research. 2019 Apr 15;25(8):2560-74.
- Ebert LM, Yu W, Gargett T, Toubia J, Kollis PM, Tea MN, Ebert BW, Bardy C, van den Hurk M, Bonder CS, Manavis J. Endothelial, pericyte and tumor cell expression in glioblastoma identifies fibroblast activation protein (FAP) as an excellent target for immunotherapy. Clinical & Translational Immunology. 2020;9(10):e1191.
- Zhou W, Ke SQ, Huang Z, Flavahan W, Fang X, Paul J, Wu L, Sloan AE, McLendon RE, Li X, Rich JN. Periostin secreted by glioblastoma stem cells recruits M2 tumour-associated macrophages and promotes malignant growth. Nature Cell Biology. 2015 Feb;17(2):170-82.
- Ceccarelli M, Barthel FP, Malta TM, Sabedot TS, Salama SR, Murray BA, Morozova O, Newton Y, Radenbaugh A, Pagnotta SM, Anjum S. Molecular profiling reveals biologically discrete subsets and pathways of progression in diffuse glioma. Cell. 2016 Jan 28;164(3):550-63.
- Cancer Genome Atlas Research Network. Comprehensive, integrative genomic analysis of diffuse lower-grade gliomas. New England Journal of Medicine. 2015 Jun 25;372(26):2481-98.
- Ravi VM, Will P, Kueckelhaus J, Sun N, Joseph K, Salié H, Vollmer L, Kuliesiute U, von Ehr J, Benotmane JK, Neidert N. Spatially resolved multi-omics deciphers bidirectional tumor-host interdependence in glioblastoma. Cancer Cell. 2022 Jun 13;40(6):639-55.
- Hao Y, Stuart T, Kowalski MH, Choudhary S, Hoffman P, Hartman A, Srivastava A, Molla G, Madad S, Fernandez-Granda C, Satija R. Dictionary learning for integrative, multimodal and scalable single-cell analysis. Nature Biotechnology. 2024 Feb;42(2):293-304.
- Robinson MD, McCarthy DJ, Smyth GK. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics. 2010 Jan 1;26(1):139-40.
- Liberzon A, Birger C, Thorvaldsdóttir H, Ghandi M, Mesirov JP, Tamayo P. The molecular signatures database hallmark gene set collection. Cell Systems. 2015 Dec 23;1(6):417-25.
- Cerami E, Gao J, Dogrusoz U, Gross BE, Sumer SO, Aksoy BA, Jacobsen A, Byrne CJ, Heuer ML, Larsson E, Antipin Y. The cBio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data. Cancer Discovery. 2012 May 1;2(5):401-4.
- Therneau TM, Grambsch PM. The cox model. In: Modeling Survival Data: Extending the Cox Model. 2000. pp. 39-77. New York, NY: Springer New York.