
July is Sarcoma Awareness Month and a good moment to talk about a cancer type that stays under-researched even as oncology continues to advance.
Sarcoma is a rare disease cancer of connective tissue - bone, fat, muscle, cartilage. It isn't a single disease but a constellation of over 70 distinct subtypes, each with its own biology. This diversity breaks the usual drug discovery playbook. Unlike common cancers such as breast or lung cancer, most sarcomas don't share one driver mutation or a clinically actionable biomarker that can be tested for or targeted.
Here are five reasons why sarcoma remains one of oncology's most challenging areas of research.
1. Extreme molecular and clinical heterogeneity
Sarcoma encompasses dozens of subtypes with distinct genetics and clinical behavior. Only a handful have simple oncogenic drivers, such as GIST (gastrointestinal stromal tumor) with its KIT and PDGFRA mutations, or synovial sarcoma with its characteristic SYT::SSX translocation. The remaining majority have complex aneuploid genomes, driven by copy-number changes and chromosomal instability rather than a single mutation that a drug can target.
The practical consequence: a target found in one subtype rarely helps another, so pan-sarcoma drugs often fail. Research has to pivot subtype by subtype, and target discovery increasingly has to look beyond mutations entirely, toward epigenetic regulators or protein complexes shared across subtypes, as in recent EZH2 and YAP pathway studies.
2. Rarity and scarce samples
Sarcomas are rare (~5–6 cases per 100,000 people per year), meaning few patients and limited tissue for research. Clinical trials struggle to accrue enough patients, and biobanks have small cohorts. As a result, statistical power is low and many subtype-specific mutations remain undiscovered. Researchers must rely on global networks and consortium biobanking to amass datasets, and may need to use advanced in silico approaches to mine TCGA/GENIE (small cohorts) for signals.
3. Complex microenvironment
Sarcomas arise in connective tissues such as bone, fat, and muscle, and may originate from different mesenchymal stem or progenitor cells. Their tumors also exist within a complex microenvironment of supporting tissue and immune cells. These interactions can influence how the tumor grows, evolves, and responds to treatment, making it difficult to identify which targets are truly driving the disease. As a result, researchers must study not only the cancer cells but also their surrounding microenvironment. Advanced 3D models and co-culture systems are helping recreate these interactions more accurately, enabling better target validation.
4. Uncertain cell of origin
Sarcomas may arise from different mesenchymal stem or progenitor cells depending on subtype, and for many subtypes the precise cell of origin is still debated. Without a clear starting point, it is harder to model disease progression accurately or to know which developmental pathways are relevant to target.
5. Limited models and translational bottlenecks
Preclinical models for sarcoma are scarce. Only a few cell lines exist for most subtypes, and patient-derived xenografts or organoids are largely limited to the more common sarcomas. Genetically engineered mouse models are also rare because of the complex genetics. Without good models, validating new targets is slow and uncertain. Patients with sarcoma are underrepresented in trials, and the small market makes large drug companies cautious. In practice this demands innovative trial designs (e.g. basket or multi-arm trials through international consortia) and underscores the need for better predictive biomarkers. Overall, these translational gaps – from limited biobanked tissue to few validated models and biomarkers – significantly slow the path from target discovery to effective sarcoma therapies.
Taken together, these translational gaps, from limited biobanked tissue to few validated models and biomarkers, significantly slow the path from target discovery to effective sarcoma therapies. The biggest takeaway from sarcoma research isn't just the shortage of data - it's the shortage of connected, evidence-backed knowledge graphs.
When every subtype tells a different biological story, progress depends less on collecting more data and more on connecting the right evidence across fragmented datasets, modalities, and publications.
This is where computational approaches are becoming increasingly important. Rather than treating genomics, transcriptomics, clinical data, imaging, and literature as separate sources of information, researchers need ways to integrate them into a unified view of disease biology.
At Elucidata, we've seen this challenge across many therapeutic areas. Our evidence-backed knowledge graphs bring together more than 30 biological and clinical data modalities, reconstruct disease–gene relationships from primary evidence, and enable researchers to apply disease-specific scoring frameworks instead of relying on generic rankings. This allows scientists to evaluate targets in the context of the biology that matters for their specific disease, even when patient cohorts are small and evidence is scattered.
We encountered a similar problem while studying neuroendocrine prostate cancer (NEPC), an aggressive cancer that is underrepresented in most public oncology datasets. Rather than relying on a single source of evidence, we built a disease-specific, evidence-backed knowledge graph that integrated diverse biological datasets to reconstruct disease biology. This approach helped uncover a novel regulatory gene-gene interaction in NEPC that conventional target discovery platforms had missed. We recently demonstrated this workflow in a webinar., showing how evidence-backed knowledge graphs can help in data-scarce diseases.
For rare diseases like sarcoma, where patient cohorts are small and evidence is fragmented, the ability to connect diverse sources of biological evidence may be just as important as generating new data.