The original MSc pipeline proved the core proposition: PubMed abstracts can be retrieved programmatically, converted into a consistent schema with a locally hosted LLM, cleaned deterministically, scored, and then analysed as structured evidence.
Project-defined scoring bands used to compare heterogeneous studies. These are analytical categories, not clinical grading thresholds.
Human evidence was the strongest subset; animal and in-vitro records were useful for biological plausibility but carried lower direct clinical relevance and lower extraction confidence.
| Evidence subset | Records | Mean evidence | Mean confidence |
|---|---|---|---|
| Human | 53 | 68.8 | 62.1 |
| Animal / preclinical | 48 | 57.2 | 33.4 |
| Mixed | 17 | 52.8 | 35.3 |
| In vitro | 8 | 45.1 | 25.6 |
Evidence and confidence are separated so that a well-designed study with an incomplete abstract is not treated the same as a clearly extracted but methodologically weak study.
Clinically mature peptide-related therapies were more likely to be represented by human studies and stronger evidence profiles. More experimental peptides were generally supported by sparser and more preclinical evidence. The important analytical result is the structure of the evidence base, not a claim that a particular compound is effective or ineffective.