<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.0 20040830//EN" "journalpublishing.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="2.0" xml:lang="en" article-type="research-article"><front><journal-meta><journal-id journal-id-type="nlm-ta">JMIR Form Res</journal-id><journal-id journal-id-type="publisher-id">formative</journal-id><journal-id journal-id-type="index">27</journal-id><journal-title>JMIR Formative Research</journal-title><abbrev-journal-title>JMIR Form Res</abbrev-journal-title><issn pub-type="epub">2561-326X</issn><publisher><publisher-name>JMIR Publications</publisher-name><publisher-loc>Toronto, Canada</publisher-loc></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">v10i1e100779</article-id><article-id pub-id-type="doi">10.2196/100779</article-id><article-categories><subj-group subj-group-type="heading"><subject>Original Paper</subject></subj-group></article-categories><title-group><article-title>TrialTriage, a Semiautonomous Prescreening Workflow for Resolving Ambiguity in Phase I Oncology Trial Eligibility: Development and Proof-of-Concept Study Using Synthetic Cases</article-title></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name name-style="western"><surname>Kern</surname><given-names>Kenneth A</given-names></name><degrees>MS, MPH, MD</degrees><xref ref-type="aff" rid="aff1">1</xref><xref ref-type="aff" rid="aff2">2</xref></contrib></contrib-group><aff id="aff1"><institution>Affiliate Faculty, Skaggs School of Pharmacy and Pharmaceutical Sciences, University of California San Diego</institution><addr-line>La Jolla</addr-line><addr-line>CA</addr-line><country>United States</country></aff><aff id="aff2"><institution>Oncology Breakthrough Advisory Group, LLC</institution><addr-line>1804 Garnet Avenue, Suite 701</addr-line><addr-line>San Diego</addr-line><addr-line>CA</addr-line><country>United States</country></aff><contrib-group><contrib contrib-type="editor"><name name-style="western"><surname>MacNeill</surname><given-names>Luke</given-names></name></contrib></contrib-group><contrib-group><contrib contrib-type="reviewer"><name name-style="western"><surname>Anandan</surname><given-names>Heber</given-names></name></contrib><contrib contrib-type="reviewer"><name name-style="western"><surname>Satasiya</surname><given-names>Kinjal</given-names></name></contrib></contrib-group><author-notes><corresp>Correspondence to Kenneth A Kern, MS, MPH, MD, Oncology Breakthrough Advisory Group, LLC, 1804 Garnet Avenue, Suite 701, San Diego, CA, 92109, United States, 1 (619) 354-1026; <email>info@oncologybreakthrough.com</email></corresp></author-notes><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>5</day><month>8</month><year>2026</year></pub-date><volume>10</volume><elocation-id>e100779</elocation-id><history><date date-type="received"><day>08</day><month>05</month><year>2026</year></date><date date-type="rev-recd"><day>30</day><month>06</month><year>2026</year></date><date date-type="accepted"><day>06</day><month>07</month><year>2026</year></date></history><copyright-statement>&#x00A9; Kenneth A Kern. Originally published in JMIR Formative Research (<ext-link ext-link-type="uri" xlink:href="https://formative.jmir.org">https://formative.jmir.org</ext-link>), 5.8.2026. </copyright-statement><copyright-year>2026</copyright-year><license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (<ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link>), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Formative Research, is properly cited. The complete bibliographic information, a link to the original publication on <ext-link ext-link-type="uri" xlink:href="https://formative.jmir.org">https://formative.jmir.org</ext-link>, as well as this copyright and license information must be included.</p></license><self-uri xlink:type="simple" xlink:href="https://formative.jmir.org/2026/1/e100779"/><abstract><sec><title>Background</title><p>Enrollment in phase I oncology trials remains low largely because potentially eligible patients are not identified and evaluated quickly enough. Current clinical trial matching systems can identify candidate patients from the electronic health record, but cases with missing or uncertain eligibility data are often routed for offline manual review. This delay impedes clarification and prolongs the final eligibility determination.</p></sec><sec><title>Objective</title><p>This study evaluated TrialTriage, a semiautonomous system built on the n8n platform and designed to resolve eligibility ambiguity during prescreening for phase I oncology trials. When eligibility information is missing or uncertain, TrialTriage emails the investigator, captures the reply, and reruns classification within the same workflow.</p></sec><sec sec-type="methods"><title>Methods</title><p>TrialTriage combined large language model&#x2013;based variable extraction from free-text clinical narratives and investigator email replies with a deterministic rule engine applying a prespecified 7-criterion protocol. Each case was classified as eligible, not eligible, or ambiguous. Ambiguous cases triggered a structured email query to the investigator, followed by reclassification after a reply. Two requests were sent at 24-hour intervals; after 48 hours without a reply, the case was referred for manual review. The system was tested on 90 synthetic patient cases generated independently by Claude Sonnet 4.6, Gemini 3.1, and Grok 4, with 30 cases per model and balanced distributions of eligible, not eligible, and ambiguous cases. Answer keys were reviewed for accuracy before system execution. Five independent reviewers classified the Claude dataset using a uniform survey form.</p></sec><sec sec-type="results"><title>Results</title><p>TrialTriage&#x2019;s classifications were 100% concordant with the author-confirmed ground truth in all 90 synthetic cases (95% CI 96.0%-100.0%). All ambiguous cases were correctly escalated to investigator query. The mean processing time was 2.3 (SD 0.5) minutes per 30-case dataset (range 1.8-2.8 min, approximately 3.5-5.5 s per case). The 5 reviewers achieved a mean accuracy of 96.7% (SD 3.3%), with a Fleiss &#x03BA; of 0.910, and required a mean of 9.8 (SD 4.8) minutes to review 30 cases. In a subset test of 6 first-pass ambiguous cases, 4 of 6 were reclassified definitively after investigator response, while 2 remained ambiguous because the replies lacked actionable information.</p></sec><sec sec-type="conclusions"><title>Conclusions</title><p>TrialTriage demonstrates the feasibility of a semiautonomous prescreening workflow in which ambiguous cases trigger an immediate investigator email query and are reclassified after reply capture with new information within the same system. The main contribution is the integration of email ambiguity resolution into the workflow rather than immediate deferral to offline manual review. Because the evaluation used synthetic cases and label definitions aligned with the same protocol rules used to design the rule engine, these findings should be interpreted as proof of concept and implementation fidelity rather than evidence of real-world clinical performance. Prospective validation using data from real-world electronic health records would be a plausible next step.</p></sec></abstract><kwd-group><kwd>clinical trials</kwd><kwd>prescreening</kwd><kwd>phase I</kwd><kwd>oncology</kwd><kwd>patient eligibility</kwd><kwd>clinical trial matching</kwd><kwd>workflow automation</kwd><kwd>large language model</kwd><kwd>LLM</kwd><kwd>natural language processing</kwd><kwd>semiautonomous AI</kwd><kwd>feasibility study</kwd><kwd>patient accrual</kwd></kwd-group></article-meta></front><body><sec id="s1" sec-type="intro"><title>Introduction</title><p>Phase I oncology trials are the initial clinical testing platform for new anticancer drugs, either alone or in combination [<xref ref-type="bibr" rid="ref1">1</xref>-<xref ref-type="bibr" rid="ref3">3</xref>]. These trials establish safety, tolerability, and preliminary dosing. Their successful completion is a prerequisite for progression to later-phase trials that support regulatory approval [<xref ref-type="bibr" rid="ref4">4</xref>]. Robust interpretation of phase I data requires the enrollment of a predefined sample size [<xref ref-type="bibr" rid="ref4">4</xref>,<xref ref-type="bibr" rid="ref5">5</xref>]. Reaching that sample size depends on identifying and enrolling patients who meet trial entry criteria through a 2-step process known as eligibility screening.</p><p>Eligibility screening for phase I oncology trials is conducted in 2 sequential steps: prescreening followed by formal screening [<xref ref-type="bibr" rid="ref5">5</xref>,<xref ref-type="bibr" rid="ref6">6</xref>]. Prescreening is an initial review of the electronic health record (EHR) or oncologist&#x2019;s patient listing against a limited set of key trial entry criteria, conducted without patient contact. Patients identified as potential candidates are then contacted, the trial is explained, and those who express interest are asked to sign an informed consent document to allow formal screening. Formal screening establishes final eligibility through detailed disease assessment, including physical examination, laboratory testing of major organ function, and radiographic evaluation of disease extent. Patients who satisfy all criteria after formal screening are enrolled [<xref ref-type="bibr" rid="ref7">7</xref>,<xref ref-type="bibr" rid="ref8">8</xref>]. Both steps are challenging in phase I oncology because of strict entry criteria, the unpredictable safety profile of novel agents, and the narrow patient populations these trials require.</p><p>Fewer than 8% of eligible patients with cancer enroll in clinical trials [<xref ref-type="bibr" rid="ref9">9</xref>-<xref ref-type="bibr" rid="ref11">11</xref>]. This low accrual rate is not primarily driven by patient refusal: between 55% and 80% of patients identified as eligible and invited agreed to enroll [<xref ref-type="bibr" rid="ref12">12</xref>-<xref ref-type="bibr" rid="ref14">14</xref>]. The dominant barrier is the inefficient identification of the smaller subset of patients who meet entry criteria from among many patients who do not [<xref ref-type="bibr" rid="ref6">6</xref>,<xref ref-type="bibr" rid="ref15">15</xref>]. Among patients who are prescreened and reach formal screening, approximately 25% fail to meet eligibility criteria and are termed &#x201C;screen failures&#x201D; [<xref ref-type="bibr" rid="ref16">16</xref>,<xref ref-type="bibr" rid="ref17">17</xref>]. Because formal screening is expensive and resource-intensive, this rate raises the question of whether better prescreening could improve the screening success rate and increase trial enrollment.</p><p>Improving low enrollment rates would have significant consequences for early-phase trial completion. Many phase I studies terminate early due to low enrollment rates and failure to meet statistically defined sample sizes. In one analysis of terminated trials, 57% closed specifically because of insufficient accrual of patients rather than drug toxicity or lack of antitumor activity [<xref ref-type="bibr" rid="ref10">10</xref>].</p><p>Phase I oncology trials typically last 2 to 3 years, with low enrollment rates as a primary contributor to delays [<xref ref-type="bibr" rid="ref5">5</xref>,<xref ref-type="bibr" rid="ref18">18</xref>]. For patients with malignant disease, these delays extend the time required for potentially effective therapies to reach them [<xref ref-type="bibr" rid="ref19">19</xref>].</p><p>Prescreening is typically performed manually by clinical research staff, who review records, query databases for missing information, and escalate uncertain cases to clinicians. These steps are labor-intensive, and a substantial fraction of patients cannot be definitively classified at this stage. Prior studies of clinical trial matching (CTM) systems report that 20% to 25% of prescreened patients cannot be definitively classified because of incomplete or ambiguous data and require manual adjudication [<xref ref-type="bibr" rid="ref20">20</xref>,<xref ref-type="bibr" rid="ref21">21</xref>]. High model performance does not resolve this problem on its own. A prospective randomized trial of 20,707 patients showed no enrollment impact despite high model performance (area under the receiver operating characteristic curve=0.91), largely because 23.2% of notifications were rejected due to a lack of clinical factors key to protocol entry [<xref ref-type="bibr" rid="ref22">22</xref>]. Delay in resolving ambiguity has direct clinical consequences: processing delays of 48 hours after imaging detects the progression of disease have resulted in 20% of patients starting alternative treatments before trial notification occurs [<xref ref-type="bibr" rid="ref23">23</xref>]. Although multiple factors contribute to accrual delays, unresolved eligibility ambiguity at the prescreening stage is a key driver of enrollment inefficiency and a documented contributor to the screen-failure rate described above [<xref ref-type="bibr" rid="ref16">16</xref>,<xref ref-type="bibr" rid="ref17">17</xref>,<xref ref-type="bibr" rid="ref24">24</xref>].</p><p>AI approaches have been developed to support patient identification through the analysis of structured and unstructured clinical data [<xref ref-type="bibr" rid="ref25">25</xref>,<xref ref-type="bibr" rid="ref26">26</xref>]. These systems, often described as CTM systems, analyze eligibility criteria and classify patients as eligible, not eligible, or ambiguous. In current CTM systems, cases classified as ambiguous are typically removed from the automated workflow and sent for manual review; information obtained during that review is generally not returned to the original classification process [<xref ref-type="bibr" rid="ref27">27</xref>]. As a result, ambiguity is not resolved within the workflow itself but is shifted into a slower manual process, increasing delays and creating additional opportunities for inconsistency or error.</p><p>To address this limitation, we designed TrialTriage, a semiautonomous workflow for prescreening in phase I oncology trials. TrialTriage classifies patients as eligible, not eligible, or ambiguous and automatically sends an email to the investigator when required information is missing or unclear. The system&#x2019;s large language model (LLM) evaluates the investigator&#x2019;s reply, extracts any substantive clarifying information, and passes that information to the rule engine for reclassification within the same workflow. This design builds on the broader principle that corrective human input can improve classification accuracy, as reported in a study in which manual feedback increased performance from 88.0% to 92.7% [<xref ref-type="bibr" rid="ref20">20</xref>]. TrialTriage extends that principle by automatically requesting feedback and clarification of ambiguous cases by email within the workflow, rather than leaving those cases for manual review outside of it.</p><p>The aim of this study was to develop and evaluate TrialTriage, a semiautonomous prescreening workflow for phase I oncology trials, whose novel component is an immediate, iterative email exchange with the principal investigator (PI). There were 2 key objectives for this study. The first was to determine whether TrialTriage could perform the following functions: correctly classify synthetic phase I oncology cases as eligible, not eligible, or ambiguous; route a listing of ambiguous cases and the data needed to classify them to an investigator through immediate email query; apply an iterative process that uses PI responses to reclassify ambiguous cases to the correct eligibility category; and retain unresolved cases as ambiguous for manual classification. The second was to compare TrialTriage&#x2019;s classification accuracy and processing time against those of human reviewers evaluating the same cases.</p><p>We hypothesized that TrialTriage would classify cases in complete agreement with the predefined ground truth and do so in less time than human reviewers, even when those reviewers were experienced in phase I trials, and that TrialTriage would request additional data from the PI and reclassify ambiguous cases accurately. Because the evaluation used synthetic cases built from the same criteria used to design the rule engine, we framed these as tests of internal validity and implementation fidelity, not of real-world clinical performance. This approach aligns with calls to reduce screen failure through automation [<xref ref-type="bibr" rid="ref28">28</xref>] and with current regulatory guidance emphasizing transparent, auditable outputs and structured human oversight for AI used in clinical trials [<xref ref-type="bibr" rid="ref29">29</xref>,<xref ref-type="bibr" rid="ref30">30</xref>].</p></sec><sec id="s2" sec-type="methods"><title>Methods</title><sec id="s2-1"><title>Study Design</title><p>This proof-of-concept study evaluated a semiautonomous workflow for prescreening synthetic patient cases against the eligibility criteria of a hypothetical phase I oncology trial. Synthetic cases were generated by LLMs to simulate patients with advanced pancreatic cancer, as described below. Pancreatic cancer was used as a representative disease to test eligibility in a hypothetical phase I oncology trial. The architecture of the workflow is not specific to any one tumor type and could be generalized to other phase I oncology trials.</p></sec><sec id="s2-2"><title>Ethical Considerations</title><p>This study did not involve human subjects and did not use actual patient data. All clinical cases were synthetic narratives generated by LLMs. Because this study involved only synthetic patient data and no living individuals underwent data collection, it did not constitute human subjects research and did not require institutional review board review or approval under the US Common Rule (45 CFR 46.102). The 5 clinicians who classified cases for the manual comparison did so as expert colleagues evaluating synthetic cases as part of a methods comparison. Since they were not research subjects, no personal data about these reviewers were recorded beyond their professional role&#x2014;physician or nonphysician clinician&#x2014;and the time taken to complete the classification task.</p></sec><sec id="s2-3"><title>System Architecture</title><p>TrialTriage is a hybrid architecture orchestrated on the n8n workflow automation platform (n8n GmbH). In n8n, individual workflow steps, called &#x201C;nodes,&#x201D; are connected through a visual interface to define a sequence of operations. The TrialTriage workflow comprises 3 functional components: intake and preprocessing, classification and iterative reclassification, and communication and output. These components are organized into 5 sequential stages: input processing, initial classification, first investigator email, ingestion of additional data for ambiguous cases, and final case reclassification and output. A simplified architecture for the TrialTriage workflow is shown in <xref ref-type="fig" rid="figure1">Figure 1</xref>.</p><p>The 5 stages are implemented using 45 deployed node instances drawn from 20 unique node types. Three of these node types incorporate LLMs: an Extract from File node, which extracts the 7 eligibility variables from free-text input; a Text Classifier node, which categorizes investigator replies as substantive or nonsubstantive; and an AI Agent node, which converts substantive reply text into structured data for reingestion by the rule engine. All 3 node types use Anthropic&#x2019;s Claude Sonnet 4.6 as the underlying language model. References to &#x201C;Agent&#x201D; elsewhere in this <italic>Methods</italic> section refer specifically to the n8n AI Agent node type.</p><p>Claude Sonnet 4.6 was selected as the model embedded within the 3 LLM-based nodes for 2 reasons. First, Claude was already integrated as a selectable model within the n8n platform, which simplified deployment and avoided the need for custom API handling. Second, the rule engine had been developed using GPT-5 (OpenAI) during Version 1 of TrialTriage, and choosing a different LLM for Version 2 provided some independence from that earlier development.</p><p>Screenshots of the workflow architecture are provided in <xref ref-type="supplementary-material" rid="app1">Multimedia Appendix 1</xref>. A complete listing of the n8n node types used in the workflow, along with their deployed instance counts, is provided in <xref ref-type="supplementary-material" rid="app2">Multimedia Appendix 2</xref>. The version of n8n used in this study retained timestamped records of each classification, email query, response ingestion, and reclassification event for up to 7 days, creating an audit trail of workflow activity. An enterprise version of n8n, not used in this report, can maintain longer-term storage.</p><fig position="float" id="figure1"><label>Figure 1.</label><caption><p>Simplified TrialTriage architecture. Patient files are received by email and validated. The large language model (LLM) extracts the 7 eligibility variables; a deterministic rule engine classifies each case as eligible, ambiguous, or not eligible. Ambiguous cases trigger an automated email to the principal investigator (PI), with one escalation if no reply is received within 24 hours. Substantive replies are reextracted by the LLM and reclassified by the rule engine. Unresolved cases exit to a recommended manual workflow. Diamond shapes indicate decision steps.</p></caption><graphic alt-version="no" mimetype="image" position="float" xlink:type="simple" xlink:href="formative_v10i1e100779_fig01.png"/></fig></sec><sec id="s2-4"><title>Data Ingestion, Validation, and Initial Classification</title><p>When a patient file is received by email, an input validation node checks both the file format and data format. Accepted file formats are XLSX, CSV, and DOCX. If either check fails, the system sends an automated email to the PI identifying the format problem and requesting resubmission. Only files and data that pass validation proceed to downstream processing. After format validation, the system&#x2019;s LLM, Claude Sonnet 4.6 (Anthropic), extracts the patient ID and eligibility variables from the free-text clinical narrative for each patient. A deterministic rule engine then assigns each case to one of 3 categories: eligible, not eligible, or ambiguous.</p></sec><sec id="s2-5"><title>Eligibility Criteria</title><p>Seven eligibility criteria, representative of those commonly used in prescreening patients with pancreatic cancer for a phase I oncology trial, were defined for this study (<xref ref-type="table" rid="table1">Table 1</xref>). Thresholds were approximate and were not intended to reproduce any specific existing trial. The system was designed to accept &#x201C;pancreatic cancer&#x201D; as synonymous with &#x201C;pancreatic adenocarcinoma&#x201D; to avoid misclassification based on terminology alone. Cases were classified as eligible only if all 7 criteria were documented in the case narrative as meeting entry requirements.</p><table-wrap id="t1" position="float"><label>Table 1.</label><caption><p>The 7 eligibility criteria used by the TrialTriage rule engine to classify synthetic phase I oncology trial cases.</p></caption><table id="table1" frame="hsides" rules="groups"><thead><tr><td align="left" valign="bottom">Criterion</td><td align="left" valign="bottom">Threshold</td></tr></thead><tbody><tr><td align="left" valign="top">Histology</td><td align="left" valign="top">Pancreatic adenocarcinoma (pancreatic cancer) confirmed by pathology within the previous 6 months</td></tr><tr><td align="left" valign="top">Age</td><td align="left" valign="top">18&#x2010;80 years, inclusive</td></tr><tr><td align="left" valign="top">ECOG<sup><xref ref-type="table-fn" rid="table1fn1">a</xref></sup> performance status</td><td align="left" valign="top">0 or 1: ECOG functional scale, 0&#x2010;4, where 0 indicates fully active and 4 indicates completely disabled and confined to bed or chair</td></tr><tr><td align="left" valign="top">AST<sup><xref ref-type="table-fn" rid="table1fn2">b</xref></sup> (SGOT<sup><xref ref-type="table-fn" rid="table1fn3">c</xref></sup>)</td><td align="left" valign="top">&#x2264;40 U/L</td></tr><tr><td align="left" valign="top">ALT<sup><xref ref-type="table-fn" rid="table1fn4">d</xref></sup> (SGPT<sup><xref ref-type="table-fn" rid="table1fn5">e</xref></sup>)</td><td align="left" valign="top">&#x2264;40 U/L</td></tr><tr><td align="left" valign="top">Total bilirubin</td><td align="left" valign="top">&#x2264;1.2 mg/dL</td></tr><tr><td align="left" valign="top">Creatinine</td><td align="left" valign="top">&#x2264;1.5 mg/dL</td></tr></tbody></table><table-wrap-foot><fn id="table1fn1"><p><sup>a</sup>ECOG: Eastern Cooperative Oncology Group.</p></fn><fn id="table1fn2"><p><sup>b</sup>AST: aspartate aminotransferase.</p></fn><fn id="table1fn3"><p><sup>c</sup>SGOT: serum glutamic-oxaloacetic transaminase.</p></fn><fn id="table1fn4"><p><sup>d</sup>ALT: alanine aminotransferase.</p></fn><fn id="table1fn5"><p><sup>e</sup>SGPT: serum glutamic-pyruvic transaminase.</p></fn></table-wrap-foot></table-wrap></sec><sec id="s2-6"><title>Classification Categories</title><p>After the system extracts the eligibility variables and the rule engine classifies each case, the system groups the results by category and sends 3 separate emails to the PI: one listing all eligible cases, one listing all not eligible cases, and one listing all ambiguous cases. When a case is classified as ambiguous, the system sends the investigator a structured email requesting the specific missing or unclear information needed for reclassification. When a substantive response is received, the workflow ingests the investigator&#x2019;s reply and reruns extraction and classification to generate an updated classification. After reclassification, the system again sends 3 category emails to the PI containing the updated eligible, not eligible, and remaining ambiguous cases. A final email asks the PI to review and approve all classifications.</p></sec><sec id="s2-7"><title>Closed-Loop Workflow Execution and Final Eligibility Decisions</title><p>The workflow executes a closed-loop sequence of data ingestion, classification, action, and reingestion without human intervention between steps. Specifically, it extracts narrative information, classifies cases through the rule engine, sends emails and schedules reminders, receives investigator responses, and reclassifies cases when new information is provided. The architecture is read-only with respect to any upstream clinical data source: the workflow does not write back to the EHR or modify original input data [<xref ref-type="bibr" rid="ref30">30</xref>]. As shown in <xref ref-type="fig" rid="figure1">Figure 1</xref>, the workflow outputs case lists for eligible, not eligible, and ambiguous categories after both the initial classification and any iterative reclassification. Final eligibility decisions for all cases remain the responsibility of the PI.</p></sec><sec id="s2-8"><title>Synthetic Dataset Generation</title><p>TrialTriage was built in 2 versions: Version 1, a limited prototype; and Version 2, the system presented in this report. Both versions used the same 7 eligibility criteria. In both versions, synthetic cases were generated using 4 LLMs, each prompted independently: ChatGPT-5.0 (OpenAI), Claude Sonnet 4.6 (Anthropic), Gemini-3.1 (Google), and Grok 4 (xAI). The same prompt structure was used across models (<xref ref-type="supplementary-material" rid="app3">Multimedia Appendix 3</xref>). Each model was instructed to generate brief but realistic synthetic case narratives of patients with pancreatic cancer, 2 to 3 sentences in length, using the 7 protocol-specific eligibility criteria described above.</p><p>Version 1, built and tested using 50 ChatGPT-generated cases, performed initial classification only. It produced eligible, not eligible, or ambiguous category emails and did not include an iterative investigator email loop or reclassification. Version 1 is not reported here, and the Version 1 synthetic cases are not included in the multimedia appendix. In Version 2 of TrialTriage, evaluated in this report, additional functions included an iterative investigator email loop, reclassification of ambiguous cases, and a substantiveness check on investigator replies.</p><p>Version 2 of TrialTriage was evaluated on cases generated independently by Claude Sonnet 4.6, Gemini-3.1, and Grok 4. Each of the 3 models generated 30 clinical scenarios in randomized order, comprising 10 eligible cases, 10 not eligible cases, and 10 ambiguous cases. Each model also produced an answer key specifying the correct classification for each case. The 90 cases were all different clinical scenarios, produced using different permutations of the 7 eligibility criteria, with no case appearing in more than 1 model&#x2019;s dataset. The synthetic case datasets were generated in February 2026. All answer keys were checked by the author to confirm the ground truth. The full case listings, with ground-truth labels and classification outputs, are provided in <xref ref-type="supplementary-material" rid="app4">Multimedia Appendix 4</xref>.</p></sec><sec id="s2-9"><title>Evaluation Procedure</title><p>The workflow included a built-in evaluation branch that permitted either a regular workflow mode, triggered by a Gmail message with a specified subject line, or a separate evaluation mode run from spreadsheet tabs stored in Google Sheets. Initial testing used the 30 Claude-generated cases. The Gemini and Grok datasets were then submitted to the system to assess concordance with independently generated inputs. TrialTriage classification of all 90 cases was performed in March 2026. TrialTriage classifications for each case are included alongside the case listings in <xref ref-type="supplementary-material" rid="app4">Multimedia Appendix 4</xref>.</p></sec><sec id="s2-10"><title>Ground Truth</title><p>Each model generated an answer key specifying the correct classification for each case. During the Grok evaluation, the author identified disagreements with the model-generated answer keys in more than 50% of cases. Investigation determined that the rule engine recognized only &#x201C;pancreatic adenocarcinoma,&#x201D; whereas some Grok-generated cases used the medically synonymous term &#x201C;pancreatic cancer.&#x201D; Both terms refer identically to cancer of the pancreas. The rule engine, which is deterministic and not a trained model, was modified to accept &#x201C;pancreatic cancer&#x201D; as equivalent. The final evaluation of all 90 cases was conducted on the modified rule engine. Classifications agreed with the author-confirmed ground truth in all 90 cases.</p><p>Because Claude served both as one of the 3 case generators and as the embedded LLM in the current TrialTriage, we checked whether TrialTriage favored cases produced by Claude by comparing its performance on cases it generated versus those generated by Gemini and Grok. Following the correction of the pancreatic cancer definition described above, classification was 100% concordant across the Claude, Gemini, and Grok cases. This single-pass evaluation was not designed to measure run-to-run variability, which remains untested and is noted as a limitation.</p></sec><sec id="s2-11"><title>Investigator Clarification for Ambiguous Cases</title><p>Cases classified as ambiguous triggered immediate email communication, with the PI requesting clarification of the specific missing or uncertain data elements. The first email listed each ambiguous case by ID along with the relevant clinical narrative and contained a &#x201C;Respond&#x201D; button. When clicked, the button opened a reply email prompting the PI to provide the specific missing or unclear data for each identified case. If no reply was received within 24 hours, a second email was sent requesting a response. If no reply was received within an additional 24 hours, a third and final email informed the PI (or, in future designs, other designated research staff) that the case was being removed from the automated classification system and needed to be reviewed through a manual process. The workflow then took no further action on classifying the case and relied on the PI to have the case evaluated by the site&#x2019;s standard manual prescreening process.</p><p>Routing the case directly from the workflow to a named secondary reviewer or coordinator was not implemented in this study and is identified in the <italic>Future Directions</italic> section as a planned enhancement. Investigator replies were evaluated by the LLM and classified as substantive or nonsubstantive. Substantive replies contained actionable clinical information relevant to one or more ambiguous eligibility criteria. Nonsubstantive replies lacked actionable clinical information, such as &#x201C;I do not have that information,&#x201D; &#x201C;I will investigate it,&#x201D; or out-of-office replies.</p></sec><sec id="s2-12"><title>Comparison of Manual Prescreening Performance to TrialTriage</title><p>Five reviewers manually classified the randomized 30-case Claude dataset using the same 7 eligibility criteria applied by TrialTriage. Reviewers were blinded to the answer key. Three reviewers were physicians with extensive drug development experience, while the remaining 2 were nonphysician clinicians with similarly extensive drug development experience. None had seen the cases prior to evaluation. The 5 reviewers completed the manual classification of the Claude dataset between March 4 and March 30, 2026. Each reviewer independently recorded classifications for all 30 cases on a separate survey sheet (<xref ref-type="supplementary-material" rid="app5">Multimedia Appendix 5</xref>), assigning each case as eligible, not eligible, or ambiguous. At the top of each survey sheet, before the case listings, instructions defined the 3 classifications and stated that any case with missing, pending, vague, or nonnumeric data should be classified as ambiguous. Average classification time per case was calculated by dividing each reviewer&#x2019;s total classification time by 30. Manual review was limited to the Claude dataset; performance against the Gemini and Grok datasets was not assessed.</p></sec><sec id="s2-13"><title>Statistical Analysis</title><p>Classification accuracy was reported with exact binomial 95% CIs. Agreement between each reviewer and the ground truth was quantified using Cohen &#x03BA; with 95% CI. Interrater agreement across all 5 reviewers was quantified using Fleiss &#x03BA;. The 95% CI was estimated by nonparametric bootstrap resampling over 5000 iterations, drawing from the table of how each reviewer classified each case. This was a post hoc analysis performed in Python 3.12.3 (Python Software Foundation) with NumPy 2.4.4 (NumFOCUS, Inc), outside the n8n workflow.</p><p>Total time required to classify the 30-case dataset was summarized descriptively. Reviewer times were reported as the mean (SD) across the 5 reviewers, while TrialTriage times were reported as the observed range across the 30 cases. TrialTriage and reviewer classification times were compared descriptively rather than inferentially because both sample sizes were small. This paper follows the TRIPOD-LLM reporting guideline, which specifies what studies that develop or evaluate LLMs in health care should report [<xref ref-type="bibr" rid="ref31">31</xref>]. A completed checklist is provided in <xref ref-type="supplementary-material" rid="app6">Checklist 1</xref>.</p></sec></sec><sec id="s3" sec-type="results"><title>Results</title><sec id="s3-1"><title>Datasets and Evaluation</title><p>TrialTriage was evaluated on the 90 synthetic cases described in <italic>Methods</italic>, comprising three 30-case datasets generated by Claude Sonnet 4.6, Gemini-3.1, and Grok 4. The Claude dataset served as the primary evaluation set. TrialTriage was applied to all 3 datasets, but the iterative investigator email clarification loop was applied only to the Claude dataset. The 5 reviewers also classified only the Claude dataset. The Gemini and Grok datasets were used to test the consistency of TrialTriage&#x2019;s performance across cases generated by different LLMs. Complete case narratives, ground-truth labels, TrialTriage classifications, and reviewer classifications are provided in <xref ref-type="supplementary-material" rid="app4">Multimedia Appendix 4</xref>.</p></sec><sec id="s3-2"><title>TrialTriage Classification Performance</title><p>TrialTriage&#x2019;s classifications were concordant with the author-confirmed ground truth in all 90 cases across the 3 evaluation datasets (90/90; exact binomial 95% CI 96.0%-100.0%), including all eligible, not eligible, and ambiguous cases. Concordance was 30/30 on the Claude dataset, 30/30 on the Gemini dataset, and 30/30 on the Grok dataset. Classification remained stable across datasets generated by different LLMs, despite differences in wording and case presentation. Because the synthetic cases were generated by LLMs prompted with the same 7 eligibility criteria from which the rule engine was constructed, this finding represents internal consistency of the workflow on prompt-faithful narratives rather than performance on the real clinical text.</p></sec><sec id="s3-3"><title>System Processing Time</title><p>TrialTriage processed each 30-case dataset in 1.8 to 2.8 minutes, corresponding to 3.5 to 5.5 seconds per case. Total processing time across all 90 cases was 5.4 to 8.4 minutes. This timing included data ingestion, LLM extraction of structured variables, and rule-engine classification. Per-dataset processing times for the Claude, Gemini, and Grok datasets all fell within this range. All 90 cases were processed without the interruption of the n8n workflow. These figures reflect automated processing time only and do not include the 24-hour escalation intervals or investigator response time required for ambiguous cases entering the email clarification loop.</p></sec><sec id="s3-4"><title>Iterative Reclassification of Ambiguous Cases</title><p>The iterative investigator feedback loop was evaluated on 6 of the 10 ambiguous cases in the Claude primary dataset, selected at random. The author served as a simulated PI and replied to the workflow&#x2019;s automated email queries; this single-responder design is acknowledged in the <italic>Limitations</italic> section. For 4 cases, the author provided additional clarifying data involving ECOG (Eastern Cooperative Oncology Group) status, confirmation of diagnosis, or laboratory values within the range. The rule engine reclassified all 4 as eligible after the LLM re-extracted variables from the reply, consistent with the data the author provided. For the remaining 2 cases, the author responded with nonsubstantive statements such as &#x201C;data not available&#x201D; or &#x201C;I don&#x2019;t have that information.&#x201D; Both cases were retained as ambiguous cases and triggered the workflow&#x2019;s recommendation for manual classification, as specified in the <italic>Methods</italic> section. All replies in this test were submitted within 24 hours, so the second-email escalation path was not triggered or evaluated.</p></sec><sec id="s3-5"><title>Manual Reviewer Performance</title><p>Reviewer accuracy on the 30-case Claude dataset ranged from 93.3% to 100.0%, with a mean of 96.7% (SD 3.3%); per-reviewer accuracy, classification time, and Cohen &#x03BA; are shown in <xref ref-type="table" rid="table2">Table 2</xref>. Interrater agreement across all 5 reviewers was almost perfect according to the Landis and Koch benchmarks [<xref ref-type="bibr" rid="ref32">32</xref>] (Fleiss &#x03BA;=0.910; bootstrap 95% CI 0.813-0.980). The mean reviewer classification time was 9.8 (SD 4.8) minutes per 30 cases, about 20 seconds per case. TrialTriage classified the same 30 cases in 1.8 to 2.8 minutes, about 4 to 6 seconds per case, with no errors. TrialTriage was therefore about 4 times faster.</p><table-wrap id="t2" position="float"><label>Table 2.</label><caption><p>Reviewer and TrialTriage performance on the 30-case Claude dataset.</p></caption><table id="table2" frame="hsides" rules="groups"><thead><tr><td align="left" valign="bottom">Reviewer</td><td align="left" valign="bottom">Role</td><td align="left" valign="bottom">Time (min)</td><td align="left" valign="bottom">Time (s)</td><td align="left" valign="bottom">Correct</td><td align="left" valign="bottom">Accuracy (%)</td><td align="left" valign="bottom">Cohen &#x03BA;<sup><xref ref-type="table-fn" rid="table2fn1">a</xref></sup> (95% CI)</td></tr></thead><tbody><tr><td align="left" valign="top">Reviewer A</td><td align="left" valign="top">MD<sup><xref ref-type="table-fn" rid="table2fn2">b</xref></sup></td><td align="left" valign="top">6.0</td><td align="left" valign="top">360</td><td align="left" valign="top">30/30</td><td align="left" valign="top">100.0</td><td align="left" valign="top">1.000<sup><xref ref-type="table-fn" rid="table2fn3">c</xref></sup></td></tr><tr><td align="left" valign="top">Reviewer B</td><td align="left" valign="top">MD</td><td align="left" valign="top">18.0</td><td align="left" valign="top">1080</td><td align="left" valign="top">30/30</td><td align="left" valign="top">100.0</td><td align="left" valign="top">1.000<sup><xref ref-type="table-fn" rid="table2fn3">c</xref></sup></td></tr><tr><td align="left" valign="top">Reviewer C</td><td align="left" valign="top">MD</td><td align="left" valign="top">6.9</td><td align="left" valign="top">415</td><td align="left" valign="top">29/30</td><td align="left" valign="top">96.7</td><td align="left" valign="top">0.950 (0.854&#x2010;1.000)</td></tr><tr><td align="left" valign="top">Reviewer D</td><td align="left" valign="top">Non-MD</td><td align="left" valign="top">10.0</td><td align="left" valign="top">600</td><td align="left" valign="top">28/30</td><td align="left" valign="top">93.3</td><td align="left" valign="top">0.900 (0.766&#x2010;1.000)</td></tr><tr><td align="left" valign="top">Reviewer E</td><td align="left" valign="top">Non-MD</td><td align="left" valign="top">8.0</td><td align="left" valign="top">480</td><td align="left" valign="top">28/30</td><td align="left" valign="top">93.3</td><td align="left" valign="top">0.900 (0.766&#x2010;1.000)</td></tr><tr><td align="left" valign="top">Mean (SD)</td><td align="left" valign="top">&#x2014;<sup><xref ref-type="table-fn" rid="table2fn4">d</xref></sup></td><td align="left" valign="top">9.8 (4.8)</td><td align="left" valign="top">587 (290)</td><td align="left" valign="top">&#x2014;</td><td align="left" valign="top">96.7 (3.3)</td><td align="left" valign="top">&#x2014;</td></tr><tr><td align="left" valign="top">TrialTriage</td><td align="left" valign="top">&#x2014;</td><td align="left" valign="top">1.8&#x2010;2.8</td><td align="left" valign="top">105&#x2010;165</td><td align="left" valign="top">30/30</td><td align="left" valign="top">100.0</td><td align="left" valign="top">1.000<sup><xref ref-type="table-fn" rid="table2fn3">c</xref></sup></td></tr></tbody></table><table-wrap-foot><fn id="table2fn1"><p><sup>a</sup>Cohen &#x03BA; is reported for each reviewer relative to the ground-truth answer key, with asymptotic 95% CI. Interrater agreement across all 5 reviewers: Fleiss &#x03BA;=0.910 (bootstrap 95% CI 0.813-0.980, 5000 resamples).</p></fn><fn id="table2fn2"><p><sup>b</sup>MD: physician.</p></fn><fn id="table2fn3"><p><sup>c</sup>The CI is not calculable when &#x03BA;=1.000, with no observed disagreement.</p></fn><fn id="table2fn4"><p><sup>d</sup>Not applicable.</p></fn></table-wrap-foot></table-wrap></sec><sec id="s3-6"><title>Manual Reviewer Errors</title><p>Five misclassifications occurred across the 150 reviewer classifications; each is listed by case, reviewer, and failure mode in <xref ref-type="table" rid="table3">Table 3</xref>. Four of the 5 errors reflected overinclusion of patients who did not meet the criteria or the premature resolution of ambiguity, and no reviewer classified an eligible case as not eligible or ambiguous. TrialTriage made no classification errors on the same dataset.</p><table-wrap id="t3" position="float"><label>Table 3.</label><caption><p>Manual reviewer classification errors in the 30-case Claude dataset<sup><xref ref-type="table-fn" rid="table3fn1">a</xref></sup>.</p></caption><table id="table3" frame="hsides" rules="groups"><thead><tr><td align="left" valign="bottom">Case ID</td><td align="left" valign="bottom">Reviewer</td><td align="left" valign="bottom">Correct class</td><td align="left" valign="bottom">Reviewer response</td><td align="left" valign="bottom">Error</td><td align="left" valign="bottom">Failure mode</td></tr></thead><tbody><tr><td align="left" valign="top">5</td><td align="left" valign="top">C</td><td align="left" valign="top">Not eligible</td><td align="left" valign="top">Eligible</td><td align="left" valign="top">Not eligible&#x2192;Eligible</td><td align="left" valign="top">Age of 83 years exceeds the upper limit of 80 years; should be classified as not eligible</td></tr><tr><td align="left" valign="top">5</td><td align="left" valign="top">D</td><td align="left" valign="top">Not eligible</td><td align="left" valign="top">Eligible</td><td align="left" valign="top">Not eligible&#x2192;Eligible</td><td align="left" valign="top">Age of 83 years exceeds the upper limit of 80 years; should be classified as not eligible</td></tr><tr><td align="left" valign="top">6</td><td align="left" valign="top">E</td><td align="left" valign="top">Ambiguous</td><td align="left" valign="top">Not eligible</td><td align="left" valign="top">Ambiguous&#x2192;Not eligible</td><td align="left" valign="top">Bilirubin absent from the lab panel; missing data should be classified as ambiguous</td></tr><tr><td align="left" valign="top">16</td><td align="left" valign="top">D</td><td align="left" valign="top">Ambiguous</td><td align="left" valign="top">Eligible</td><td align="left" valign="top">Ambiguous&#x2192;Eligible</td><td align="left" valign="top">No ECOG<sup><xref ref-type="table-fn" rid="table3fn2">b</xref></sup> numeric status, narrative only; should be classified as ambiguous</td></tr><tr><td align="left" valign="top">18</td><td align="left" valign="top">E</td><td align="left" valign="top">Not eligible</td><td align="left" valign="top">Eligible</td><td align="left" valign="top">Not eligible&#x2192;Eligible</td><td align="left" valign="top">SGPT<sup><xref ref-type="table-fn" rid="table3fn3">c</xref></sup> 57 U/L exceeds ULN<sup><xref ref-type="table-fn" rid="table3fn4">d</xref></sup> of 40 U/L; should be classified as not eligible</td></tr></tbody></table><table-wrap-foot><fn id="table3fn1"><p><sup>a</sup>Error column shows the misclassification as correct class&#x2192;reviewer response.</p></fn><fn id="table3fn2"><p><sup>b</sup>ECOG: Eastern Cooperative Oncology Group.</p></fn><fn id="table3fn3"><p><sup>c</sup>SGPT: serum glutamic-pyruvic transaminase.</p></fn><fn id="table3fn4"><p><sup>d</sup>ULN: upper limit of normal.</p></fn></table-wrap-foot></table-wrap></sec></sec><sec id="s4" sec-type="discussion"><title>Discussion</title><sec id="s4-1"><title>Principal Findings</title><p>TrialTriage met both objectives of this proof-of-concept evaluation. The first objective was to determine whether TrialTriage could classify synthetic phase I oncology cases, route ambiguous cases to the investigator, reclassify them based on the PI&#x2019;s response, and retain unresolved cases for manual classification. TrialTriage successfully classified every synthetic case in agreement with the predefined ground truth across datasets generated by 3 different LLMs, categorizing each as eligible, not eligible, or ambiguous. It routed every ambiguous case through an immediate, structured email query, reclassified the cases that received substantive replies, and retained the remaining cases as Ambiguous for manual classification. The second objective was to compare TrialTriage&#x2019;s classification accuracy and processing time against those of human reviewers evaluating the same cases. TrialTriage classified cases faster than the reviewers and without classification errors, whereas the reviewers&#x2019; errors involved either overinclusion or premature resolution of ambiguous cases.</p></sec><sec id="s4-2"><title>Comparison With Prior Work</title><p>Existing CTM systems and TrialTriage can be compared based on 2 points: what they do with the cases they cannot classify and how accurately they classify. Current systems focus their automated workflow primarily on identifying eligible patients, and they divert ambiguous cases for offline, manual resolution, which is often delayed [<xref ref-type="bibr" rid="ref33">33</xref>-<xref ref-type="bibr" rid="ref35">35</xref>]. When clinical observations are the only source for an eligibility variable, or when variables related to the past medical history are absent or unclear, automated CTMs cannot resolve the case definitively, and the case is routed for manual review. Manual review of such cases is time-consuming and resource-intensive [<xref ref-type="bibr" rid="ref6">6</xref>]. TrialTriage operates downstream of these systems, prescreening identified candidates against a specific trial&#x2019;s criteria and resolving ambiguity within the same workflow.</p><p>We surveyed published descriptions of CTM and prescreening systems, including Tempus Tapp [<xref ref-type="bibr" rid="ref36">36</xref>], TrialMatchAI [<xref ref-type="bibr" rid="ref37">37</xref>], OncoLLM [<xref ref-type="bibr" rid="ref38">38</xref>], Watson CTM [<xref ref-type="bibr" rid="ref21">21</xref>], TrialGPT [<xref ref-type="bibr" rid="ref39">39</xref>], and a natural language processing&#x2013;based screen failure prediction model [<xref ref-type="bibr" rid="ref19">19</xref>]. None of these systems describe an automated mechanism for resolving ambiguous cases within the workflow. Several do not address ambiguous cases at all, leaving it unclear whether such cases are excluded from analysis, defaulted to one classification, or routed for offline manual review.</p><p>In a cohort of 102 patients with non&#x2013;small cell lung cancer, IBM Watson CTM had a median processing time of 15.5 seconds per case (range 7.2&#x2010;37.8 s) and achieved 97% agreement with human review [<xref ref-type="bibr" rid="ref40">40</xref>]. Nevertheless, the workflow required 8088 manual actions to enter data or obtain clinician interpretation for eligibility, and Watson CTM did not query for missing information during the automated workflow. Reported accuracies for CTMs reflect performance largely on cases with sufficient information already available because the data-incomplete cases are excluded from the accuracy denominator [<xref ref-type="bibr" rid="ref38">38</xref>].</p></sec><sec id="s4-3"><title>Implications</title><p>This proof-of-concept study holds several implications for the use of semiautonomous systems in real-world eligibility determinations. This type of workflow may address enrollment inefficiencies because prescreening and formal enrollment screening are interdependent. Inaccurate prescreening advances potential study participants who subsequently fail to enroll due to the more detailed formal screening process [<xref ref-type="bibr" rid="ref27">27</xref>]. McKane et al [<xref ref-type="bibr" rid="ref16">16</xref>] reported a 24.6% screen-failure rate in phase I trials, which rose 3-fold to 78.2% at a center that included inaccurate prescreening classification as part of the overall screen-failure rate. These failed screening efforts are also costly ($25,000 per patient) and resource-intensive (up to 8.5 h of staff time per patient) [<xref ref-type="bibr" rid="ref41">41</xref>,<xref ref-type="bibr" rid="ref42">42</xref>]. Overall, phase I trials commonly require accrual times 5-fold longer than originally planned because of the logistical obstacles created by complex eligibility criteria, such as biomarker and biopsy data [<xref ref-type="bibr" rid="ref5">5</xref>,<xref ref-type="bibr" rid="ref43">43</xref>,<xref ref-type="bibr" rid="ref44">44</xref>]. Automated workflows, such as TrialTriage, may help to reduce costs and resource time by classifying cases accurately and quickly and by resolving ambiguous cases through an immediate email query to the investigator, although these benefits remain to be confirmed with real clinical data.</p><p>Bias can enter TrialTriage at the points where LLMs are used: extracting variables from the case, judging whether an investigator&#x2019;s reply is substantive, and parsing substantive reply text into the data fields read by the rule engine. Several features of the architecture guard against errors at these points. Eligibility decisions are made by a deterministic rule engine applying threshold values, so the model does not rely on its own judgment. Investigator replies pass through a substantiveness check, and noninformative replies trigger an email recommending manual prescreening. The eligibility output is sent to the PI or research staff for final sign-off, and a timestamped audit trail records all classifications, queries, replies, and reclassifications. To prevent contamination of the EHR, TrialTriage reads only the source record and never writes back to it.</p><p>Some subjectivity and bias are likely to remain because some eligibility parameters, such as ECOG performance status (a graded measure of a patient&#x2019;s ability to carry out daily activities), rely strongly on clinical judgment and cannot be classified by objective measures alone.</p></sec><sec id="s4-4"><title>Limitations</title><p>This study has several limitations, including the use of synthetic rather than real clinical cases, the limited independence of the test cases from the rules used to build the rule engine, small sample sizes, the limited set of eligibility criteria tested, and the absence of real-world data governance.</p><sec id="s4-4-1"><title>Synthetic Cases</title><p>The evaluation used synthetic case narratives rather than EHR-derived records. Synthetic cases are useful for testing workflow logic and rule implementation; however, because of their inherent simplicity, they cannot exactly reproduce the semantic and syntactic complexity of unstructured, real clinical documentation [<xref ref-type="bibr" rid="ref38">38</xref>,<xref ref-type="bibr" rid="ref45">45</xref>]. Actual medical records often contain inconsistencies, redundancies, conflicting documentation, and highly variable narrative quality, all of which complicate the extraction of clinically meaningful information [<xref ref-type="bibr" rid="ref46">46</xref>-<xref ref-type="bibr" rid="ref49">49</xref>]. Whether TrialTriage&#x2019;s performance on synthetic cases would be maintained on real-world EHR narratives remains untested. Because the synthetic cases were not designed to reflect any patient demographic mix, the evaluation also cannot demonstrate how TrialTriage performs across different patient groups.</p></sec><sec id="s4-4-2"><title>Test-Case Independence</title><p>The Version 2 test cases generated by Claude, Gemini, and Grok were built from variations on the same 7 eligibility criteria used to construct the rule engine during Version 1 development. The use of 3 different LLMs broadened the diversity of how cases were described but did not create a fully independent evaluation set. Performance against an external standard independent of the rule-based development process remains to be established.</p></sec><sec id="s4-4-3"><title>Sample Size</title><p>The ambiguity-resolution analysis was limited to 6 randomly selected ambiguous cases from the Claude dataset, with the author responding to the system&#x2019;s email queries as a simulated PI. The workflow handled those cases correctly; however, the sample is too small to support strong conclusions about how the system would perform with real investigator replies that vary in clarity, completeness, and timing. Whether real PIs or qualified study personnel would respond quickly enough to preserve the workflow&#x2019;s time advantage remains untested. The reviewer comparison involved 5 clinical reviewers evaluating one 30-case dataset. That sample was sufficient to reveal general error patterns, particularly overinclusion and premature classification of ambiguous cases as eligible or not eligible, but it was not large enough to establish stable benchmarks for manual prescreening performance or to identify which case features most reliably produce reviewer error. The reviewer comparison should therefore be understood as hypothesis-generating rather than definitive.</p></sec><sec id="s4-4-4"><title>Data Governance</title><p>Because the study used only synthetic cases and deliberately did not include protected health information, it cannot show how actual patient data would be governed within the workflow. Before using a semiautonomous system like TrialTriage in a phase I clinical trial, procedures for handling uploaded patient data must be established to ensure confidentiality. These procedures would specify where identifiable patient data enter the workflow, which steps process or store data, how long data are retained, how data are encrypted in storage and in transit, and how the investigator email channel is protected.</p></sec><sec id="s4-4-5"><title>Criteria Scope</title><p>The workflow was tested on 7 representative criteria rather than the full complexity of a phase I protocol. Formal trial screening commonly involves more criteria, including biomarker requirements, disease extent, prior treatment history, organ function, and many other features derived from physical examination, laboratory data, and imaging. Typical phase I oncology protocols have upward of 60 eligibility criteria that must be met during formal screening [<xref ref-type="bibr" rid="ref40">40</xref>]. Whether the same approach remains effective as criterion count and complexity expand will require further study.</p></sec></sec><sec id="s4-5"><title>Future Directions</title><p>This proof-of-concept evaluation leaves 3 questions about the TrialTriage architecture unresolved, which could form the basis for future development.</p><sec id="s4-5-1"><title>Extraction Layer</title><p>The next stage of development would test TrialTriage on EHR-derived cases, with the goals of characterizing extraction performance on clinician-generated text and determining whether immediate, iterative email communication can shorten the time to final prescreening classification under real-world conditions. Operational end points for such testing include time to final prescreening decision, frequency of ambiguity resolution within a defined time window after the first email request, percentage of workload shifted away from manual review, prescreening failure rate (the number of patients referred for formal screening who were subsequently rejected), and screen-failure rate (the percentage of consented patients who were never dosed).</p></sec><sec id="s4-5-2"><title>Human Factors</title><p>The performance of the human in the loop should be measured by the time taken to respond to the iterative emails and by the quality of the data supplied in reply. Implementations could direct email not only to the PI but also to designated research staff or coordinators, allowing a comparison of which responders provide the greatest gain in efficiency. TrialTriage must ultimately be tested in its integrated form, combining the semiautonomous classification workflow and the human-in-the-loop email response, to determine the overall accuracy and timing of the complete system.</p></sec><sec id="s4-5-3"><title>Scalability</title><p>How the rule engine behaves as eligibility criteria increase, perhaps by a factor of 10 or more, is unknown. Testing performance with substantially more eligibility criteria must also include the human responders, whose replies are essential to the iterative process.</p></sec></sec><sec id="s4-6"><title>Conclusions</title><p>TrialTriage demonstrated the feasibility of resolving prescreening ambiguity through immediate investigator email communication and iterative reclassification within the same automated workflow. If validated on EHR-derived cases and complex phase I trial criteria, the TrialTriage architecture could reduce manual prescreening workload and improve prescreening accuracy in early-phase oncology trials.</p></sec></sec></body><back><ack><p>The author thanks Ryan Nolan, BSEE, for his implementation of the n8n workflow according to the author&#x2019;s specifications, along with the physicians and clinicians who contributed their clinical judgment to the manual classification survey. Draw.io was used to create <xref ref-type="fig" rid="figure1">Figure 1</xref>. Generative AI tools were used in the preparation of this manuscript. The literature search was carried out using SciSpace and Consensus. Claude Sonnet 4.6 assisted with text editing and generated <xref ref-type="table" rid="table1">Tables 1</xref><xref ref-type="table" rid="table2"/>-<xref ref-type="table" rid="table3">3</xref>. The author reviewed and verified all AI-generated content, including literature references, and takes full responsibility for the manuscript. No AI tool is listed as an author.</p></ack><notes><sec><title>Funding</title><p>The author declared that no financial support was received for this work. The platform vendor, n8n, did not sponsor this work, provide financial support, or compensate the consultant. The author used a standard self-paid subscription and personally funded all costs.</p></sec><sec><title>Data Availability</title><p>All data generated and analyzed during this study are included in the manuscript and its multimedia appendices.</p></sec></notes><fn-group><fn fn-type="con"><p>Conceptualization: KAK</p><p>Data curation: KAK</p><p>Formal analysis: KAK</p><p>Investigation: KAK</p><p>Methodology: KAK</p><p>Validation: KAK</p><p>Visualization: KAK</p><p>Writing &#x2013; original draft: KAK</p><p>Writing &#x2013; review and editing: KAK</p></fn><fn fn-type="conflict"><p>None declared.</p></fn></fn-group><glossary><title>Abbreviations</title><def-list><def-item><term id="abb1">CTM</term><def><p>clinical trial matching</p></def></def-item><def-item><term id="abb2">ECOG</term><def><p>Eastern Cooperative Oncology Group</p></def></def-item><def-item><term id="abb3">EHR</term><def><p>electronic health record</p></def></def-item><def-item><term id="abb4">LLM</term><def><p>large language model</p></def></def-item><def-item><term id="abb5">PI</term><def><p>principal investigator</p></def></def-item></def-list></glossary><ref-list><title>References</title><ref id="ref1"><label>1</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Araujo</surname><given-names>D</given-names> </name><name name-style="western"><surname>Greystoke</surname><given-names>A</given-names> </name><name name-style="western"><surname>Bates</surname><given-names>S</given-names> </name><etal/></person-group><article-title>Oncology phase I trial design and conduct: time for a change-MDICT Guidelines 2022</article-title><source>Ann Oncol</source><year>2023</year><month>01</month><volume>34</volume><issue>1</issue><fpage>48</fpage><lpage>60</lpage><pub-id pub-id-type="doi">10.1016/j.annonc.2022.09.158</pub-id><pub-id pub-id-type="medline">36182023</pub-id></nlm-citation></ref><ref id="ref2"><label>2</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Alotaibi</surname><given-names>H</given-names> </name><name name-style="western"><surname>Anis</surname><given-names>AM</given-names> </name><name name-style="western"><surname>Alloghbi</surname><given-names>A</given-names> </name><name name-style="western"><surname>Alshammari</surname><given-names>K</given-names> </name></person-group><article-title>Oncology early-phase clinical trials in the Middle East and North Africa: a review of the current status, challenges, opportunities, and future directions</article-title><source>J Immunother Precis Oncol</source><year>2024</year><month>08</month><volume>7</volume><issue>3</issue><fpage>178</fpage><lpage>189</lpage><pub-id pub-id-type="doi">10.36401/JIPO-23-25</pub-id><pub-id pub-id-type="medline">39219998</pub-id></nlm-citation></ref><ref id="ref3"><label>3</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Weber</surname><given-names>JS</given-names> </name><name name-style="western"><surname>Levit</surname><given-names>LA</given-names> </name><name name-style="western"><surname>Adamson</surname><given-names>PC</given-names> </name><etal/></person-group><article-title>American Society of Clinical Oncology policy statement update: the critical role of phase I trials in cancer research and treatment</article-title><source>J Clin Oncol</source><year>2015</year><month>01</month><day>20</day><volume>33</volume><issue>3</issue><fpage>278</fpage><lpage>284</lpage><pub-id pub-id-type="doi">10.1200/JCO.2014.58.2635</pub-id></nlm-citation></ref><ref id="ref4"><label>4</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Rubin</surname><given-names>EH</given-names> </name><name name-style="western"><surname>Gilliland</surname><given-names>DG</given-names> </name></person-group><article-title>Drug development and clinical trials&#x2014;the path to an approved cancer drug</article-title><source>Nat Rev Clin Oncol</source><year>2012</year><month>02</month><day>28</day><volume>9</volume><issue>4</issue><fpage>215</fpage><lpage>222</lpage><pub-id pub-id-type="doi">10.1038/nrclinonc.2012.22</pub-id><pub-id pub-id-type="medline">22371130</pub-id></nlm-citation></ref><ref id="ref5"><label>5</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Massett</surname><given-names>HA</given-names> </name><name name-style="western"><surname>Mishkin</surname><given-names>G</given-names> </name><name name-style="western"><surname>Rubinstein</surname><given-names>L</given-names> </name><etal/></person-group><article-title>Challenges facing early phase trials sponsored by the National Cancer Institute: an analysis of corrective action plans to improve accrual</article-title><source>Clin Cancer Res</source><year>2016</year><month>11</month><day>15</day><volume>22</volume><issue>22</issue><fpage>5408</fpage><lpage>5416</lpage><pub-id pub-id-type="doi">10.1158/1078-0432.CCR-16-0338</pub-id><pub-id pub-id-type="medline">27401246</pub-id></nlm-citation></ref><ref id="ref6"><label>6</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Denicoff</surname><given-names>AM</given-names> </name><name name-style="western"><surname>Ivy</surname><given-names>SP</given-names> </name><name name-style="western"><surname>Tamashiro</surname><given-names>TT</given-names> </name><etal/></person-group><article-title>Implementing modernized eligibility criteria in US National Cancer Institute clinical trials</article-title><source>J Natl Cancer Inst</source><year>2022</year><month>11</month><day>14</day><volume>114</volume><issue>11</issue><fpage>1437</fpage><lpage>1440</lpage><pub-id pub-id-type="doi">10.1093/jnci/djac152</pub-id><pub-id pub-id-type="medline">36047830</pub-id></nlm-citation></ref><ref id="ref7"><label>7</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Ni</surname><given-names>Y</given-names> </name><name name-style="western"><surname>Bermudez</surname><given-names>M</given-names> </name><name name-style="western"><surname>Kennebeck</surname><given-names>S</given-names> </name><name name-style="western"><surname>Liddy-Hicks</surname><given-names>S</given-names> </name><name name-style="western"><surname>Dexheimer</surname><given-names>J</given-names> </name></person-group><article-title>A real-time automated patient screening system for clinical trials eligibility in an emergency department: design and evaluation</article-title><source>JMIR Med Inform</source><year>2019</year><month>07</month><day>24</day><volume>7</volume><issue>3</issue><fpage>e14185</fpage><pub-id pub-id-type="doi">10.2196/14185</pub-id><pub-id pub-id-type="medline">31342909</pub-id></nlm-citation></ref><ref id="ref8"><label>8</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Xiang</surname><given-names>JJ</given-names> </name><name name-style="western"><surname>Roy</surname><given-names>A</given-names> </name><name name-style="western"><surname>Summers</surname><given-names>C</given-names> </name><etal/></person-group><article-title>Brief report: implementation of a universal prescreening protocol to increase recruitment to lung cancer studies at a Veterans Affairs Cancer Center</article-title><source>JTO Clin Res Rep</source><year>2022</year><month>07</month><volume>3</volume><issue>7</issue><fpage>100357</fpage><pub-id pub-id-type="doi">10.1016/j.jtocrr.2022.100357</pub-id><pub-id pub-id-type="medline">35815320</pub-id></nlm-citation></ref><ref id="ref9"><label>9</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Carlisle</surname><given-names>B</given-names> </name><name name-style="western"><surname>Kimmelman</surname><given-names>J</given-names> </name><name name-style="western"><surname>Ramsay</surname><given-names>T</given-names> </name><name name-style="western"><surname>MacKinnon</surname><given-names>N</given-names> </name></person-group><article-title>Unsuccessful trial accrual and human subjects protections: an empirical analysis of recently closed trials</article-title><source>Clin Trials</source><year>2015</year><month>02</month><volume>12</volume><issue>1</issue><fpage>77</fpage><lpage>83</lpage><pub-id pub-id-type="doi">10.1177/1740774514558307</pub-id><pub-id pub-id-type="medline">25475878</pub-id></nlm-citation></ref><ref id="ref10"><label>10</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Williams</surname><given-names>RJ</given-names> </name><name name-style="western"><surname>Tse</surname><given-names>T</given-names> </name><name name-style="western"><surname>DiPiazza</surname><given-names>K</given-names> </name><name name-style="western"><surname>Zarin</surname><given-names>DA</given-names> </name></person-group><article-title>Terminated trials in the ClinicalTrials.gov results database: evaluation of availability of primary outcome data and reasons for termination</article-title><source>PLoS ONE</source><year>2015</year><volume>10</volume><issue>5</issue><fpage>e0127242</fpage><pub-id pub-id-type="doi">10.1371/journal.pone.0127242</pub-id><pub-id pub-id-type="medline">26011295</pub-id></nlm-citation></ref><ref id="ref11"><label>11</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Wornow</surname><given-names>M</given-names> </name><name name-style="western"><surname>Lozano</surname><given-names>A</given-names> </name><name name-style="western"><surname>Dash</surname><given-names>D</given-names> </name><name name-style="western"><surname>Jindal</surname><given-names>J</given-names> </name><name name-style="western"><surname>Mahaffey</surname><given-names>KW</given-names> </name><name name-style="western"><surname>Shah</surname><given-names>NH</given-names> </name></person-group><article-title>Zero-shot clinical trial patient matching with LLMs</article-title><source>NEJM AI</source><year>2025</year><month>01</month><volume>2</volume><issue>1</issue><fpage>AIcs2400360</fpage><pub-id pub-id-type="doi">10.1056/AIcs2400360</pub-id></nlm-citation></ref><ref id="ref12"><label>12</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Unger</surname><given-names>JM</given-names> </name><name name-style="western"><surname>Xiao</surname><given-names>H</given-names> </name><name name-style="western"><surname>Vaidya</surname><given-names>R</given-names> </name><etal/></person-group><article-title>The cost of doing business: drug costs in federally sponsored cancer clinical trials</article-title><source>JCO Oncol Pract</source><year>2024</year><month>10</month><volume>20</volume><fpage>10</fpage><pub-id pub-id-type="doi">10.1200/OP.2024.20.10_suppl.10</pub-id></nlm-citation></ref><ref id="ref13"><label>13</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Unger</surname><given-names>JM</given-names> </name><name name-style="western"><surname>Hershman</surname><given-names>DL</given-names> </name><name name-style="western"><surname>Till</surname><given-names>C</given-names> </name><etal/></person-group><article-title>&#x201C;When offered to participate&#x201D;: a systematic review and meta-analysis of patient agreement to participate in cancer clinical trials</article-title><source>J Natl Cancer Inst</source><year>2021</year><month>03</month><day>1</day><volume>113</volume><issue>3</issue><fpage>244</fpage><lpage>257</lpage><pub-id pub-id-type="doi">10.1093/jnci/djaa155</pub-id><pub-id pub-id-type="medline">33022716</pub-id></nlm-citation></ref><ref id="ref14"><label>14</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Canou&#x00EF;-Poitrine</surname><given-names>F</given-names> </name><name name-style="western"><surname>Li&#x00E8;vre</surname><given-names>A</given-names> </name><name name-style="western"><surname>Dayde</surname><given-names>F</given-names> </name><etal/></person-group><article-title>Inclusion of older patients with cancer in clinical trials: the SAGE prospective multicenter cohort survey</article-title><source>Oncologist</source><year>2019</year><month>12</month><volume>24</volume><issue>12</issue><fpage>e1351</fpage><lpage>e1359</lpage><pub-id pub-id-type="doi">10.1634/theoncologist.2019-0166</pub-id><pub-id pub-id-type="medline">31324663</pub-id></nlm-citation></ref><ref id="ref15"><label>15</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Calaprice-Whitty</surname><given-names>D</given-names> </name><name name-style="western"><surname>Galil</surname><given-names>K</given-names> </name><name name-style="western"><surname>Salloum</surname><given-names>W</given-names> </name><name name-style="western"><surname>Zariv</surname><given-names>A</given-names> </name><name name-style="western"><surname>Jimenez</surname><given-names>B</given-names> </name></person-group><article-title>Improving clinical trial participant prescreening with artificial intelligence (AI): a comparison of the results of AI-assisted vs standard methods in 3 oncology trials</article-title><source>Ther Innov Regul Sci</source><year>2020</year><month>01</month><volume>54</volume><issue>1</issue><fpage>69</fpage><lpage>74</lpage><pub-id pub-id-type="doi">10.1007/s43441-019-00030-4</pub-id><pub-id pub-id-type="medline">32008227</pub-id></nlm-citation></ref><ref id="ref16"><label>16</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Mckane</surname><given-names>A</given-names> </name><name name-style="western"><surname>Sima</surname><given-names>C</given-names> </name><name name-style="western"><surname>Ramanathan</surname><given-names>RK</given-names> </name><etal/></person-group><article-title>Determinants of patient screen failures in phase 1 clinical trials</article-title><source>Invest New Drugs</source><year>2013</year><month>06</month><volume>31</volume><issue>3</issue><fpage>774</fpage><lpage>779</lpage><pub-id pub-id-type="doi">10.1007/s10637-012-9894-7</pub-id><pub-id pub-id-type="medline">23135779</pub-id></nlm-citation></ref><ref id="ref17"><label>17</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Campillo-Gimenez</surname><given-names>B</given-names> </name><name name-style="western"><surname>Buscail</surname><given-names>C</given-names> </name><name name-style="western"><surname>Zekri</surname><given-names>O</given-names> </name><etal/></person-group><article-title>Improving the pre-screening of eligible patients in order to increase enrollment in cancer clinical trials</article-title><source>Trials</source><year>2015</year><month>01</month><day>16</day><volume>16</volume><issue>1</issue><fpage>15</fpage><pub-id pub-id-type="doi">10.1186/s13063-014-0535-7</pub-id><pub-id pub-id-type="medline">25592642</pub-id></nlm-citation></ref><ref id="ref18"><label>18</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Frankel</surname><given-names>PH</given-names> </name><name name-style="western"><surname>Chung</surname><given-names>V</given-names> </name><name name-style="western"><surname>Tuscano</surname><given-names>J</given-names> </name><etal/></person-group><article-title>Model of a queuing approach for patient accrual in phase 1 oncology studies</article-title><source>JAMA Netw Open</source><year>2020</year><month>05</month><day>1</day><volume>3</volume><issue>5</issue><fpage>e204787</fpage><pub-id pub-id-type="doi">10.1001/jamanetworkopen.2020.4787</pub-id><pub-id pub-id-type="medline">32401317</pub-id></nlm-citation></ref><ref id="ref19"><label>19</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Delorme</surname><given-names>J</given-names> </name><name name-style="western"><surname>Charvet</surname><given-names>V</given-names> </name><name name-style="western"><surname>Wartelle</surname><given-names>M</given-names> </name><etal/></person-group><article-title>Natural language processing for patient selection in phase I or II oncology clinical trials</article-title><source>JCO Clin Cancer Inform</source><year>2021</year><month>06</month><volume>5</volume><issue>5</issue><fpage>709</fpage><lpage>718</lpage><pub-id pub-id-type="doi">10.1200/CCI.21.00003</pub-id><pub-id pub-id-type="medline">34197179</pub-id></nlm-citation></ref><ref id="ref20"><label>20</label><nlm-citation citation-type="other"><person-group person-group-type="author"><name name-style="western"><surname>Ferber</surname><given-names>D</given-names> </name><name name-style="western"><surname>Hilgers</surname><given-names>L</given-names> </name><name name-style="western"><surname>Wiest</surname><given-names>IC</given-names> </name><etal/></person-group><article-title>End-to-end clinical trial matching with large language models</article-title><source>arXiv</source><comment>Preprint posted online on  Jul 18, 2024</comment><pub-id pub-id-type="doi">10.48550/arXiv.2407.13463</pub-id></nlm-citation></ref><ref id="ref21"><label>21</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Haddad</surname><given-names>T</given-names> </name><name name-style="western"><surname>Helgeson</surname><given-names>JM</given-names> </name><name name-style="western"><surname>Pomerleau</surname><given-names>KE</given-names> </name><etal/></person-group><article-title>Accuracy of an artificial intelligence system for cancer clinical trial eligibility screening: retrospective pilot study</article-title><source>JMIR Med Inform</source><year>2021</year><month>03</month><day>26</day><volume>9</volume><issue>3</issue><fpage>e27767</fpage><pub-id pub-id-type="doi">10.2196/27767</pub-id><pub-id pub-id-type="medline">33769304</pub-id></nlm-citation></ref><ref id="ref22"><label>22</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Mazor</surname><given-names>T</given-names> </name><name name-style="western"><surname>Farhat</surname><given-names>KS</given-names> </name><name name-style="western"><surname>Trukhanov</surname><given-names>P</given-names> </name><etal/></person-group><article-title>Clinical trial notifications triggered by artificial intelligence&#x2013;detected cancer progression: a randomized trial</article-title><source>JAMA Netw Open</source><year>2025</year><month>04</month><day>1</day><volume>8</volume><issue>4</issue><fpage>e252013</fpage><pub-id pub-id-type="doi">10.1001/jamanetworkopen.2025.2013</pub-id><pub-id pub-id-type="medline">40257799</pub-id></nlm-citation></ref><ref id="ref23"><label>23</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Kehl</surname><given-names>KL</given-names> </name><name name-style="western"><surname>Mazor</surname><given-names>T</given-names> </name><name name-style="western"><surname>Trukhanov</surname><given-names>P</given-names> </name><etal/></person-group><article-title>Identifying oncology clinical trial candidates using artificial intelligence predictions of treatment change: a pilot implementation study</article-title><source>JCO Precis Oncol</source><year>2024</year><month>03</month><volume>8</volume><issue>8</issue><fpage>e2300507</fpage><pub-id pub-id-type="doi">10.1200/PO.23.00507</pub-id><pub-id pub-id-type="medline">38513166</pub-id></nlm-citation></ref><ref id="ref24"><label>24</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Ozaki</surname><given-names>H</given-names> </name><name name-style="western"><surname>Miyawaki</surname><given-names>E</given-names> </name><name name-style="western"><surname>Miyazaki</surname><given-names>N</given-names> </name><etal/></person-group><article-title>Patient characteristics related to screening failure in phase I trials</article-title><source>BMC Cancer</source><year>2025</year><month>12</month><day>4</day><volume>26</volume><issue>1</issue><fpage>69</fpage><pub-id pub-id-type="doi">10.1186/s12885-025-15311-5</pub-id><pub-id pub-id-type="medline">41339809</pub-id></nlm-citation></ref><ref id="ref25"><label>25</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Nievas</surname><given-names>M</given-names> </name><name name-style="western"><surname>Basu</surname><given-names>A</given-names> </name><name name-style="western"><surname>Wang</surname><given-names>Y</given-names> </name><name name-style="western"><surname>Singh</surname><given-names>H</given-names> </name></person-group><article-title>Distilling large language models for matching patients to clinical trials</article-title><source>J Am Med Inform Assoc</source><year>2024</year><month>09</month><day>1</day><volume>31</volume><issue>9</issue><fpage>1953</fpage><lpage>1963</lpage><pub-id pub-id-type="doi">10.1093/jamia/ocae073</pub-id><pub-id pub-id-type="medline">38641416</pub-id></nlm-citation></ref><ref id="ref26"><label>26</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Meystre</surname><given-names>SM</given-names> </name><name name-style="western"><surname>Heider</surname><given-names>PM</given-names> </name><name name-style="western"><surname>Cates</surname><given-names>A</given-names> </name><etal/></person-group><article-title>Piloting an automated clinical trial eligibility surveillance and provider alert system based on artificial intelligence and standard data models</article-title><source>BMC Med Res Methodol</source><year>2023</year><month>04</month><day>11</day><volume>23</volume><issue>1</issue><fpage>88</fpage><pub-id pub-id-type="doi">10.1186/s12874-023-01916-6</pub-id><pub-id pub-id-type="medline">37041475</pub-id></nlm-citation></ref><ref id="ref27"><label>27</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Wu</surname><given-names>J</given-names> </name><name name-style="western"><surname>Yakubov</surname><given-names>A</given-names> </name><name name-style="western"><surname>Abdul-Hay</surname><given-names>M</given-names> </name><etal/></person-group><article-title>Prescreening to increase therapeutic oncology trial enrollment at the largest public hospital in the United States</article-title><source>JCO Oncol Pract</source><year>2022</year><month>04</month><volume>18</volume><issue>4</issue><fpage>e620</fpage><lpage>e625</lpage><pub-id pub-id-type="doi">10.1200/OP.21.00629</pub-id><pub-id pub-id-type="medline">34748371</pub-id></nlm-citation></ref><ref id="ref28"><label>28</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>La Rosa</surname><given-names>A</given-names> </name><name name-style="western"><surname>Vaterkowski</surname><given-names>M</given-names> </name><name name-style="western"><surname>Cuggia</surname><given-names>M</given-names> </name><etal/></person-group><article-title>&#x201C;The truth is, we must miss some&#x201D;: a qualitative study of the patient eligibility screening process, and automation perspectives, for cancer clinical trials</article-title><source>Cancer Med</source><year>2024</year><month>12</month><volume>13</volume><issue>23</issue><fpage>e70466</fpage><pub-id pub-id-type="doi">10.1002/cam4.70466</pub-id><pub-id pub-id-type="medline">39624972</pub-id></nlm-citation></ref><ref id="ref29"><label>29</label><nlm-citation citation-type="report"><article-title>Considerations for the use of artificial intelligence to support regulatory decision-making for drug and biological products: guidance for industry and other interested parties</article-title><year>2025</year><access-date>2026-07-10</access-date><publisher-name>U.S. Food and Drug Administration</publisher-name><comment><ext-link ext-link-type="uri" xlink:href="https://www.fda.gov/media/184830/download">https://www.fda.gov/media/184830/download</ext-link></comment></nlm-citation></ref><ref id="ref30"><label>30</label><nlm-citation citation-type="report"><article-title>Guiding principles of good AI practice in drug development</article-title><year>2026</year><access-date>2026-07-10</access-date><publisher-name>European Medicines Agency</publisher-name><comment><ext-link ext-link-type="uri" xlink:href="https://www.ema.europa.eu/en/documents/other/guiding-principles-good-ai-practice-drug-development_en.pdf">https://www.ema.europa.eu/en/documents/other/guiding-principles-good-ai-practice-drug-development_en.pdf</ext-link></comment></nlm-citation></ref><ref id="ref31"><label>31</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Gallifant</surname><given-names>J</given-names> </name><name name-style="western"><surname>Afshar</surname><given-names>M</given-names> </name><name name-style="western"><surname>Ameen</surname><given-names>S</given-names> </name><etal/></person-group><article-title>The TRIPOD-LLM reporting guideline for studies using large language models</article-title><source>Nat Med</source><year>2025</year><month>01</month><volume>31</volume><issue>1</issue><fpage>60</fpage><lpage>69</lpage><pub-id pub-id-type="doi">10.1038/s41591-024-03425-5</pub-id><pub-id pub-id-type="medline">39779929</pub-id></nlm-citation></ref><ref id="ref32"><label>32</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Landis</surname><given-names>JR</given-names> </name><name name-style="western"><surname>Koch</surname><given-names>GG</given-names> </name></person-group><article-title>The measurement of observer agreement for categorical data</article-title><source>Biometrics</source><year>1977</year><month>03</month><volume>33</volume><issue>1</issue><fpage>159</fpage><lpage>174</lpage><pub-id pub-id-type="doi">10.2307/2529310</pub-id><pub-id pub-id-type="medline">843571</pub-id></nlm-citation></ref><ref id="ref33"><label>33</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Embi</surname><given-names>PJ</given-names> </name><name name-style="western"><surname>Jain</surname><given-names>A</given-names> </name><name name-style="western"><surname>Clark</surname><given-names>J</given-names> </name><name name-style="western"><surname>Bizjack</surname><given-names>S</given-names> </name><name name-style="western"><surname>Hornung</surname><given-names>R</given-names> </name><name name-style="western"><surname>Harris</surname><given-names>CM</given-names> </name></person-group><article-title>Effect of a clinical trial alert system on physician participation in trial recruitment</article-title><source>Arch Intern Med</source><year>2005</year><month>10</month><day>24</day><volume>165</volume><issue>19</issue><fpage>2272</fpage><lpage>2277</lpage><pub-id pub-id-type="doi">10.1001/archinte.165.19.2272</pub-id><pub-id pub-id-type="medline">16246994</pub-id></nlm-citation></ref><ref id="ref34"><label>34</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>K&#x00F6;pcke</surname><given-names>F</given-names> </name><name name-style="western"><surname>Prokosch</surname><given-names>HU</given-names> </name></person-group><article-title>Employing computers for the recruitment into clinical trials: a comprehensive systematic review</article-title><source>J Med Internet Res</source><year>2014</year><month>07</month><day>1</day><volume>16</volume><issue>7</issue><fpage>e161</fpage><pub-id pub-id-type="doi">10.2196/jmir.3446</pub-id><pub-id pub-id-type="medline">24985568</pub-id></nlm-citation></ref><ref id="ref35"><label>35</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Ni</surname><given-names>Y</given-names> </name><name name-style="western"><surname>Wright</surname><given-names>J</given-names> </name><name name-style="western"><surname>Perentesis</surname><given-names>J</given-names> </name><etal/></person-group><article-title>Increasing the efficiency of trial-patient matching: automated clinical trial eligibility pre-screening for pediatric oncology patients</article-title><source>BMC Med Inform Decis Mak</source><year>2015</year><month>04</month><day>14</day><volume>15</volume><issue>1</issue><fpage>28</fpage><pub-id pub-id-type="doi">10.1186/s12911-015-0149-3</pub-id><pub-id pub-id-type="medline">25881112</pub-id></nlm-citation></ref><ref id="ref36"><label>36</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Mallahan</surname><given-names>S</given-names> </name><name name-style="western"><surname>Gajra</surname><given-names>A</given-names> </name><name name-style="western"><surname>Blau</surname><given-names>S</given-names> </name><etal/></person-group><article-title>Optimizing clinical trial subject selection: insights from Exigent Research Network and the Tempus AI TIME Program collaboration</article-title><source>AI Precis Oncol</source><year>2024</year><month>12</month><day>1</day><volume>1</volume><issue>6</issue><fpage>306</fpage><lpage>314</lpage><pub-id pub-id-type="doi">10.1089/aipo.2024.0030</pub-id></nlm-citation></ref><ref id="ref37"><label>37</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Abdallah</surname><given-names>M</given-names> </name><name name-style="western"><surname>Nakken</surname><given-names>S</given-names> </name><name name-style="western"><surname>Georges</surname><given-names>M</given-names> </name><etal/></person-group><article-title>TrialMatchAI: an end-to-end AI-powered clinical trial recommendation system to streamline patient-to-trial matching</article-title><source>Nat Commun</source><year>2026</year><month>03</month><day>25</day><volume>17</volume><issue>1</issue><fpage>4472</fpage><pub-id pub-id-type="doi">10.1038/s41467-026-70509-w</pub-id><pub-id pub-id-type="medline">41876500</pub-id></nlm-citation></ref><ref id="ref38"><label>38</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Gupta</surname><given-names>S</given-names> </name><name name-style="western"><surname>Basu</surname><given-names>A</given-names> </name><name name-style="western"><surname>Nievas</surname><given-names>M</given-names> </name><etal/></person-group><article-title>PRISM: patient records interpretation for semantic clinical trial matching system using large language models</article-title><source>NPJ Digit Med</source><year>2024</year><month>10</month><day>28</day><volume>7</volume><issue>1</issue><fpage>305</fpage><pub-id pub-id-type="doi">10.1038/s41746-024-01274-7</pub-id><pub-id pub-id-type="medline">39468259</pub-id></nlm-citation></ref><ref id="ref39"><label>39</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Jin</surname><given-names>Q</given-names> </name><name name-style="western"><surname>Wang</surname><given-names>Z</given-names> </name><name name-style="western"><surname>Floudas</surname><given-names>CS</given-names> </name><etal/></person-group><article-title>Matching patients to clinical trials with large language models</article-title><source>Nat Commun</source><year>2024</year><month>11</month><day>18</day><volume>15</volume><issue>1</issue><fpage>9074</fpage><pub-id pub-id-type="doi">10.1038/s41467-024-53081-z</pub-id><pub-id pub-id-type="medline">39562565</pub-id></nlm-citation></ref><ref id="ref40"><label>40</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Alexander</surname><given-names>M</given-names> </name><name name-style="western"><surname>Solomon</surname><given-names>B</given-names> </name><name name-style="western"><surname>Ball</surname><given-names>DL</given-names> </name><etal/></person-group><article-title>Evaluation of an artificial intelligence clinical trial matching system in Australian lung cancer patients</article-title><source>JAMIA Open</source><year>2020</year><month>07</month><volume>3</volume><issue>2</issue><fpage>209</fpage><lpage>215</lpage><pub-id pub-id-type="doi">10.1093/jamiaopen/ooaa002</pub-id><pub-id pub-id-type="medline">32734161</pub-id></nlm-citation></ref><ref id="ref41"><label>41</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Hernando-Calvo</surname><given-names>A</given-names> </name><name name-style="western"><surname>Nguyen</surname><given-names>P</given-names> </name><name name-style="western"><surname>Bedard</surname><given-names>PL</given-names> </name><etal/></person-group><article-title>Impact on costs and outcomes of multi-gene panel testing for advanced solid malignancies: a cost-consequence analysis using linked administrative data</article-title><source>EClinicalMedicine</source><year>2024</year><month>03</month><volume>69</volume><fpage>102443</fpage><pub-id pub-id-type="doi">10.1016/j.eclinm.2024.102443</pub-id><pub-id pub-id-type="medline">38380071</pub-id></nlm-citation></ref><ref id="ref42"><label>42</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Penberthy</surname><given-names>LT</given-names> </name><name name-style="western"><surname>Dahman</surname><given-names>BA</given-names> </name><name name-style="western"><surname>Petkov</surname><given-names>VI</given-names> </name><name name-style="western"><surname>DeShazo</surname><given-names>JP</given-names> </name></person-group><article-title>Effort required in eligibility screening for clinical trials</article-title><source>J Oncol Pract</source><year>2012</year><month>11</month><volume>8</volume><issue>6</issue><fpage>365</fpage><lpage>370</lpage><pub-id pub-id-type="doi">10.1200/JOP.2012.000646</pub-id><pub-id pub-id-type="medline">23598846</pub-id></nlm-citation></ref><ref id="ref43"><label>43</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Dienstmann</surname><given-names>R</given-names> </name><name name-style="western"><surname>Garralda</surname><given-names>E</given-names> </name><name name-style="western"><surname>Aguilar</surname><given-names>S</given-names> </name><etal/></person-group><article-title>Evolving landscape of molecular prescreening strategies for oncology early clinical trials</article-title><source>JCO Precis Oncol</source><year>2020</year><volume>4</volume><fpage>PO.19.00398</fpage><pub-id pub-id-type="doi">10.1200/PO.19.00398</pub-id><pub-id pub-id-type="medline">32923891</pub-id></nlm-citation></ref><ref id="ref44"><label>44</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Middleton</surname><given-names>G</given-names> </name><name name-style="western"><surname>Fletcher</surname><given-names>P</given-names> </name><name name-style="western"><surname>Popat</surname><given-names>S</given-names> </name><etal/></person-group><article-title>The National Lung Matrix Trial of personalized therapy in lung cancer</article-title><source>Nature</source><year>2020</year><month>07</month><volume>583</volume><issue>7818</issue><fpage>807</fpage><lpage>812</lpage><pub-id pub-id-type="doi">10.1038/s41586-020-2481-8</pub-id><pub-id pub-id-type="medline">32669708</pub-id></nlm-citation></ref><ref id="ref45"><label>45</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Beigi</surname><given-names>M</given-names> </name><name name-style="western"><surname>Shafquat</surname><given-names>A</given-names> </name><name name-style="western"><surname>Mezey</surname><given-names>J</given-names> </name><name name-style="western"><surname>Aptekar</surname><given-names>J</given-names> </name></person-group><article-title>Simulants: synthetic clinical trial data via subject-level privacy-preserving synthesis</article-title><source>AMIA Annu Symp Proc</source><year>2023</year><volume>2022</volume><fpage>231</fpage><lpage>240</lpage><pub-id pub-id-type="medline">37128411</pub-id></nlm-citation></ref><ref id="ref46"><label>46</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Hripcsak</surname><given-names>G</given-names> </name><name name-style="western"><surname>Albers</surname><given-names>DJ</given-names> </name></person-group><article-title>Correlating electronic health record concepts with healthcare process events</article-title><source>J Am Med Inform Assoc</source><year>2013</year><month>12</month><volume>20</volume><issue>e2</issue><fpage>e311</fpage><lpage>8</lpage><pub-id pub-id-type="doi">10.1136/amiajnl-2013-001922</pub-id><pub-id pub-id-type="medline">23975625</pub-id></nlm-citation></ref><ref id="ref47"><label>47</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Weiskopf</surname><given-names>NG</given-names> </name><name name-style="western"><surname>Weng</surname><given-names>C</given-names> </name></person-group><article-title>Methods and dimensions of electronic health record data quality assessment: enabling reuse for clinical research</article-title><source>J Am Med Inform Assoc</source><year>2013</year><month>01</month><day>1</day><volume>20</volume><issue>1</issue><fpage>144</fpage><lpage>151</lpage><pub-id pub-id-type="doi">10.1136/amiajnl-2011-000681</pub-id><pub-id pub-id-type="medline">22733976</pub-id></nlm-citation></ref><ref id="ref48"><label>48</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Kahn</surname><given-names>MG</given-names> </name><name name-style="western"><surname>Callahan</surname><given-names>TJ</given-names> </name><name name-style="western"><surname>Barnard</surname><given-names>J</given-names> </name><etal/></person-group><article-title>A harmonized data quality assessment terminology and framework for the secondary use of electronic health record data</article-title><source>EGEMS (Wash DC)</source><year>2016</year><volume>4</volume><issue>1</issue><fpage>1244</fpage><pub-id pub-id-type="doi">10.13063/2327-9214.1244</pub-id><pub-id pub-id-type="medline">27713905</pub-id></nlm-citation></ref><ref id="ref49"><label>49</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Botsis</surname><given-names>T</given-names> </name><name name-style="western"><surname>Hartvigsen</surname><given-names>G</given-names> </name><name name-style="western"><surname>Chen</surname><given-names>F</given-names> </name><name name-style="western"><surname>Weng</surname><given-names>C</given-names> </name></person-group><article-title>Secondary use of EHR: data quality issues and informatics opportunities</article-title><source>Summit Transl Bioinform</source><year>2010</year><month>03</month><day>1</day><volume>2010</volume><fpage>1</fpage><lpage>5</lpage><pub-id pub-id-type="medline">21347133</pub-id></nlm-citation></ref></ref-list><app-group><supplementary-material id="app1"><label>Multimedia Appendix 1</label><p>Screenshots of the deployed TrialTriage n8n workflow, showing the complete workflow canvas and detailed views of each functional section.</p><media xlink:href="formative_v10i1e100779_app1.docx" xlink:title="DOCX File, 1521 KB"/></supplementary-material><supplementary-material id="app2"><label>Multimedia Appendix 2</label><p>Table of n8n node types used in the TrialTriage workflow, with the function and number of deployed instances of each.</p><media xlink:href="formative_v10i1e100779_app2.docx" xlink:title="DOCX File, 17 KB"/></supplementary-material><supplementary-material id="app3"><label>Multimedia Appendix 3</label><p>Prompt template used for generation of the synthetic pancreatic cancer case set across the 3 large language models reported in this study (Claude Sonnet 4.6, Gemini 3.1, and Grok 4).</p><media xlink:href="formative_v10i1e100779_app3.docx" xlink:title="DOCX File, 18 KB"/></supplementary-material><supplementary-material id="app4"><label>Multimedia Appendix 4</label><p>Datasets generated and analyzed during the study, organized by source (large language model or human reviewer). Each tab contains the synthetic case classifications, the ground truth, and the corresponding TrialTriage workflow outputs.</p><media xlink:href="formative_v10i1e100779_app4.xlsx" xlink:title="XLSX File, 40 KB"/></supplementary-material><supplementary-material id="app5"><label>Multimedia Appendix 5</label><p>Survey instrument used to collect case classifications and completion times from human reviewers for comparison against the TrialTriage workflow.</p><media xlink:href="formative_v10i1e100779_app5.docx" xlink:title="DOCX File, 25 KB"/></supplementary-material><supplementary-material id="app6"><label>Checklist 1</label><p>TRIPOD-LLM checklist.</p><media xlink:href="formative_v10i1e100779_app6.docx" xlink:title="DOCX File, 34 KB"/></supplementary-material></app-group></back></article>