A Provenance Audit of Formula and Component Evidence in Herbal Medicine Knowledge Bases
Abstract
Given an herbal formula and a protein target, we study how much evidence reported for the formula can be explained by evidence associated with its individual herbs. This question defines a supervised task in which links between formulas and targets serve as labels for evidence prediction and knowledge graph completion. Component evidence can reach those labels directly when a member herb is already linked to the target. It can also enter through computational studies that begin with component databases and later contribute formula claims to the knowledge base. Publication frequency can create another source of predictable structure because frequently studied formulas, herbs, and targets accumulate more reported links. Our outcome records the reported evidence status of a formula and target in a specified context and time. Formula identity depends on composition, dose, processing, and preparation. We represent each formula as a versioned specification that records its materials, source species, medicinal parts, processing, quantities, preparation, dosage form, route, and source. Source specifications and tested interventions receive separate records. We assign an executability tier from a material list to an observed executable intervention. These tiers determine whether a record supports analysis based on composition or a context matched comparison. For each formula and target, we compare the two summaries of component evidence below. | Summary | Definition and purpose | | --- | --- | | Broad union | Includes all target evidence associated with any member herb. It measures coverage available from a component lookup. | | Matched union | Includes member herb evidence matched on preparation, dose, route, biological context, assay direction, and publication time. It measures context and time compatible component support. | HERB 2.0 [1] catalogs 6,743 formulae, of which 159 (2.4%) have curated target or disease evidence. The full release also contains 6,644 curated references and 2,231 high throughput experiments. We organize supervision as individual formula, target, and reference claims so that shared sources remain visible. Provenance annotation operates at the level of individual claims because computational predictions and experimental measurements can coexist in one paper. Each record links its formula instance, target, direction, biological context, reference, evidence level, and derivation path. Claim type and evidence level are recorded separately. Evidence levels cover direct engagement, functional dependence, target activity, molecular abundance, pathway association, and computational prediction. This separation preserves the different support attached to claims from the same study. The initial audit samples 100 claims. A meaningful subset receives independent labels from two annotators and expert adjudication. For computational claims, we trace whether their derivation begins with member materials and databases of component target associations before applying enrichment or docking. Source overlap and lineage confidence determine the circularity status. Evaluation tests whether component evidence provides predictive signal beyond frequency and literature attention. Baselines cover the broad component union, target frequency, formula and herb publication frequency, composition, component evidence, and literature attention. The benchmark uses positive and unlabeled learning across alternative candidate universes and sampling schemes, including popularity matched sampling. Aliases, near identical compositions, modified prescriptions, and products are clustered into formula families before splitting. Primary splits separate both formula families and cited references. Temporal evaluation restricts all features to evidence available at prediction time. Classification and ranking metrics are accompanied by confidence intervals from formula cluster bootstrap. We report broad and matched union coverage separately. The central provenance comparison evaluates all claims and the set obtained after excluding circular claims. Together, these analyses provide a provenance grounded benchmark for measuring how much reported formula evidence is supported by component knowledge and how much remains after accounting for context, time, and research attention. Reference. [1] K. Gao et al. HERB 2.0: an updated TCM database integrating clinical and experimental evidence. Nucleic Acids Research, 53(D1):D1404–D1414, 2025.