Hieu Nghiem, MSa,b, Hemanth Reddy Singareddy, BSc, Zhuqi Miao, PhDb,d,*, Jivan Lamichhane, MDe , Abdulaziz Ahmed,PhDf , Johnson Thomas, PhDa , Dursun Delen, PhDb,d,g, William Paiva, PhDb

a Department of Computer Science, Oklahoma State University, Stillwater, OK, 74078, USA
b Center for Health Systems Innovation, Oklahoma State University, Tulsa, OK, 74119, USA
c Department of Computer Science, The State University of New York at New Paltz, New Paltz, NY,
12561, USA
d Department of Management Science and Information Systems, Oklahoma State University,
Stillwater, OK, 74106, USA
e Department of Medicine, The State University of New York Upstate Medical University, Syracuse,
NY, 13210, USA
f Department of Health Services Administration, University of Alabama at Birmingham,
Birmingham, AL, 35233, USA
g Department of Industrial Engineering, Faculty of Engineering and Natural Sciences, Istinye
University, Sariyer/Istanbul 34396, Turkey

Abstract
Objective: Develop a cost-effective, large language model (LLM)-based pipeline for
automatically extracting Review of Systems (ROS) entities from clinical notes.
Materials and Methods: The pipeline extracts ROS sections using SecTag, followed by few-shot LLMs to identify ROS entity spans, their positive/negative status, and associated body systems. We implemented the pipeline using open-source LLMs (Mistral, Llama, Gemma) and ChatGPT. The evaluation was conducted on 36 general medicine notes containing 341 annotated ROS entities.
Results: When integrating ChatGPT, the pipeline achieved the lowest error rates in
detecting ROS entity spans and their corresponding statuses/systems (28.2% and 14.5%, respectively). Open-source LLMs enable local, cost-efficient execution of the pipeline while delivering promising performance with similarly low error rates (span: 30.5–36.7%; status/system: 24.3–27.3%).


Discussion and Conclusion: Our pipeline offers a scalable and locally deployable solution to reduce ROS documentation burden. Open-source LLMs present a viable alternative to commercial models in resource-limited healthcare environments.


Keywords: review of systems, clinical note, natural language processing, large language model, open-source, LangChain pipeline.


Introduction
The Review of Systems (ROS) is an inventory of signs and symptoms organized by body systems.1,2 It plays an important role in evidence-based assessment by guiding clinicians in prioritizing specific systems for further evaluation during the objective examination.3,4 Additionally, the comprehensive overview provided by ROS helps uncover underlying conditions and differentiate between diseases with overlapping symptoms, ultimately enhancing diagnostic accuracy.5–8
ROS data are typically collected through verbal screenings or questionnaires at clinical encounters and are often included as a separate section within clinical notes.9 A variety of approaches have been used to document ROS, including free-text entry, dictation, and checklists supplemented with brief free-text descriptions.2,10 However, documenting, reviewing, and analyzing free text within clinical notes constitute a persistent challenge to healthcare professionals. According to the literature, U.S. physicians spend 17% to 43% of their time interacting with electronic health record (EHR) systems for documentationrelated
tasks.11–13 This significant administrative burden detracts from their primary
responsibility of providing patient care, ultimately reducing productivity and diminishing career satisfaction.14 Despite recent efforts by the American Medical Association (AMA) and the Centers for Medicare & Medicaid Services (CMS) to streamline medical documentation,15,16 documenting medically appropriate patient history, such as history of present illness (HPI) and ROS, remain entrenched in current screening and documentation workflows, continuing to contribute to the documentation burden for health professionals.17 Natural language processing (NLP), empowered by large language models (LLMs), offers a promising avenue for analyzing free-text clinical notes and extracting valuable information.18–22 However, studies on applying NLP to ROS remains limited in the literature, and recent work used fine-tuned BERT models rather than trending LLMs.23
BERT-based technologies typically require significant engineering effort, such as taskspecific input/output processing and fine-tuning on labeled datasets, to perform well on specific tasks. In contrast, LLMs are more flexible and can be adapted to a wide range of tasks in an out-of-the-box manner using natural language prompts that are easily understood by lay users. Additionally, LLMs generally have significantly more parameters, enabling them to encode a broader range of knowledge within the model itself.
This study proposes an easy-to-implement pipeline that integrates open-source LLMs for ROS recognition on local, cost-effective devices. This approach is better suited to use cases in medical documentation, offering a more accessible and scalable solution.


Methods
Clinical Notes We employed Medical Transcription Sample Reports and Examples (MTSamples) as our data source. MTSamples is a publicly available repository of de-identified clinical notes that is widely used in the medical informatics research community.24 The notes in MTSamples are organized into sections, each with a header followed by free text, as illustrated in Figure 1(A).

Figure 1. An example of MTSamples notes: (A) The sample note in plain text format; (B) The sample note with annotations.


We selected 36 general medicine notes from MTSamples that included a ROS section. The sample comprised of 19 consultation notes, 2 discharge summaries, 4 ED notes, 6 progress notes, 4 history and physical notes, and 1 urgent care note. Of these, 12 notes contained simplistic ROS sections lacking details, using phrases such as “Noncontributory,” “Otherwise negative,” or “As per the HPI.” Including these types of notes enhances the dataset’s diversity and allows for a better evaluation of our pipeline in handling varied styles in the ROS entity recognition task.
From the selected notes, we annotated 𝑛 = 341 entities in the ROS sections, including diseases, symptoms, and body systems. Each entity was labeled with its status (positive or negative) and assigned to one of the 14 standard body systems: constitutional, eyes, ENT (ears, nose, mouth, and throat), cardiovascular, respiratory, hematologic/lymphatic, gastrointestinal, genitourinary, musculoskeletal, integumentary, neurological, psychiatric, endocrine, and allergic/immunologic.25 The annotation process resulted in a total of 341
spans and 682 labels, as exemplified in Figure 1(B).


Pipeline Design


Our pipeline, illustrated in Figure 2, comprises three consecutive steps: Segmentation, ROS Entity Recognition, and Body System Classification. Each of these steps are further described below.

Figure 2: Overview of the proposed pipeline


 Segmentation: This step segments clinical notes and extracts the ROS section to ensure that downstream steps focus specifically on this section. Specifically, we used the SecTag section header terminology26,27 to identify the start and end boundaries of the ROS section and extract the corresponding text.
 ROS Entity Recognition: This step employed few-shot LLMs28,29 to extract ROS
entities from the isolated ROS section. The LLM also evaluated the positive or negative status of each entity based on the context. The few-shot examples were provided not only to guide the recognition process but also to illustrate the desired JSON output format.
 Body System Classification: This step classifies the extracted ROS entities by body system. We employed few-short LLMs to implement the classification. Since many diseases and symptoms are self-indicative of their corresponding systems, the few-shot examples primarily serve to guide the model toward generating consistent disease–system pairs, a format easy to parse using regular expressions.30 The resulting systems are then added to the JSON outputs from the previous step. Note that we separated this classification step from the prior Disease and Symptom Recognition step to avoid overly lengthy prompts, which in our preliminary experiments increased confusion and reduced model accuracy.
Once the body system was identified, we performed a Valid System check to address an issue of arbitrary phrase extraction, which frequently occurred when processing simplistic ROS sections. Specifically, many of these phrases did not represent valid ROS entities. If LLMs classified them under a category that did not match any of the 14 recognized systems, the entities were discarded.
Implementation and Evaluation To develop a simple and cost-effective pipeline, we adopted medium- or small-sized opensource LLMs that can run on consumer/workstation-grade GPUs. Specifically, we employed Mistral-small 3.1 (24B parameters, size of 15GB),31 Llama 3.2 (3B, 2GB),32 and Gemma 3 (27B, 17GB).33 These models were integrated into our pipeline, with each step automated using LangChain.34 We evaluated the pipeline on an NVIDIA RTX 3090 GPU with 24GB VRAM. To benchmark performance, we compared the results of these opensource LLM-based pipelines with those obtained by manually executing the same pipeline using the ChatGPT web application (powered by GPT-4o). The full set of LLM prompts used in the proposed pipeline is provided in the Supplementary Material. The LangChain source code for the pipeline, along with annotated notes, are available on GitHub (see the Data and Code Availability Statements). To evaluate ROS entity extraction, we used the following metrics:
 Exact Match (𝐸): The span of the ROS detection exactly matches the annotation.
 Relaxed Match (𝑅 ): The span of the ROS detection does not precisely match but
overlaps with the annotation.
 Under Detection (𝑈): Fail to detect an annotated ROS entity.
 Over Detection (𝑂): Misidentifying unannotated text as a ROS entity.35
The number of span errors that require manual correction can be estimated as:
𝑆𝑝𝑎𝑛 𝐸𝑟𝑟𝑜𝑟𝑠 = 𝑅 + 𝑈 + 𝑂 We additionally counted the correct status detections (denoted as 𝑇􀮾 ,𝑇􀯋 for exact matches and relaxed matches, respectively) and the correct system classifications (𝑌􀮾 ,𝑌􀯋 ). Label errors requiring manual correction can be indicated by the number of incorrect statuses or systems in both exact and relaxed matches, plus the counts of statuses and systems associated with under- and over-detections. This discussion leads to the formula below:
𝐿𝑎𝑏𝑒𝑙 𝐸𝑟𝑟𝑜𝑟𝑠 = 2(𝐸 + 𝑅 + 𝑈 + 𝑂) − (𝑇􀮾 + 𝑇􀯋 + 𝑌􀮾 + 𝑌􀯋 )
We used span errors and label errors as the primary performance metrics to estimate the potential manual effort the proposed pipeline can save.
Results Table 1 summarizes the performance of the pipeline based on each LLM. The ChatGPTbased pipeline achieved the lowest error rate at 28.2%. Pipelines using Mistral, Llama, and Gemma exhibited slightly higher error rates of 30.5%, 34.4%, and 36.7%, respectively. As a state-of-the-art LLM, ChatGPT has been benchmarked to outperform most open-source LLMs.36 Therefore, it is expected that the pipeline leveraging ChatGPT will achieve better performance than those using other LLMs. However, the performance gain is relatively modest, ranging from 2.3% to 8.5%. Figure 3 shows that all models achieved high accuracy (>93.0%) in detecting ROS statuses for exactly/relaxedly matched entities. For body system classification, ChatGPT maintained strong accuracy at 95.2%, while open-source LLMs also performed well, with accuracies ranging from 83.1% to 88.1%. The strong performance in status detection and system classification contributed to the low label error rates, ranging from 14.5% to 27.3%, as presented in Table 1.


Table 1: ROS recognition performance of the pipeline by LLMs: Counts of matches and errors.

Note: The span error rates represent the number of span errors relative to the total number of spans (341), indicating the percentage of spans that require correction. Similarly, the label error rates are calculated based on the total number of label errors relative to the total number of labels (682).

Figure 3. Accuracy of ROS status detection and body system classification for exactly/relaxedly matched entities across different models. Status detection accuracy = (𝑇􀮾 + 𝑇􀯋 )/(𝐸 + 𝑅); system classification accuracy = (𝑌􀮾 + 𝑌􀯋 )/(𝐸 + 𝑅).

In addition, we observed several common patterns related to span errors:
 Many relaxed matches resulted from entity rephrasing. For example, “fevers” was
detected as “fever,” and “concerns about her skin” was matched as “skin concerns.”
Another frequent source of relaxed matches was entity splitting. For instance, “muscle or joint pain” was detected as two separate entities: “muscle pain” and “joint pain.”
 One clinical note had a ROS section that was only partially captured by SecTag,
resulting in 15 under-detected entities.
 Simplistic ROS narratives could lead to hallucinations, a known issue with LLMs.37 They cannot be addressed by our Valid System check, resulting in a small amount of over detections.
Furthermore, it is worth noting that Llama utilizes only 2GB of GPU memory, in contrast to Mistral’s 15GB and Gemma’s 17GB. As a result, Llama executes significantly faster than the other two open-source LLMs. The efficiency is especially valuable when computational resources are limited in pipeline deployment.


Discussion
Documentation has long been a burden for health professionals. Although the AMA and CMS have simplified requirements, specifically by removing the need to document detailed history and examination, clinical documentation still consumes a significant amount of time in practice. Much of this time is spent extracting key information to complete claims templates.
This study presents a novel LLM-based pipeline that is capable of automatically batchprocessing clinical notes to extract ROS entities, determine their status, and classify them into corresponding body systems. The pipeline can be used with moderately sized, opensource LLMs on lightweight devices that can be easily deployed within the local infrastructure of healthcare institutions, thereby protecting patient data and offering a costeffective solution. Our evaluation demonstrated the pipeline’s generalizability and efficiency. When integrated into the pipeline, all evaluated LLMs succeeded to reduce the ROS entity recognition workload remarkably compared to a fully manual process: Only 28.8%-36.7% of the spans and 14.5%–27.3% labels still required manual correction.
Our results also showed the variability in performance that may arise from the choice of underlying LLMs. Key considerations when selecting an LLM include accuracy, execution efficiency, device requirements, and usage cost. ChatGPT offers the highest accuracy but incurs consistent usage costs. Llama is a lightweight yet effective option that is well-suited for resource-constrained environments. When resources permit, Mistral also presents a strong alternative.
Similar to the burden of recognizing and documenting ROS, documenting HPI and
Physical Examinations poses an even greater challenge for health professionals.17 Our current work on ROS extraction can be extended to identify HPI-associated signs and symptoms, helping to streamline the HPI documentation process. Another interesting technical enhancement for future work would be mitigating the LLM’s rephrasing tendency, allowing relaxed matches to be refined into exact matches.
Limitations: The clinical notes used in this study were from a single data source and
focused on general medicine. Note formats and content can vary across institutions and specialties. Evaluating the pipeline on a broader range of samples will be instrumental in enhancing its generalizability and performance. The pipeline’s performance is primarily measured by error rates, which help estimate how many ROS entities may require manual correction. However, this is only an approximation. The actual time savings achievable in real-world settings should be evaluated through practical trials.


Conclusion
An open-source LLM-based pipeline for automating ROS entity recognition is developed and evaluated in this study. The pipeline is easy and cost effective to implement and can integrate various open-source LLMs to achieve comparable accuracy as state-of-the-art ChatGPT. Future work will be focused on accuracy improvement and adaptation to HPI entity recognition.


Author Contributions
Hieu Nghiem: Methodology, Software, Investigation, Formal analysis, and Writing–
Review & Editing. Hemanth Reddy Singareddy: Data Curation, Investigation, Software and Formal analysis. Zhuqi Miao: Conceptualization, Project Administration, Data Curation, Methodology, Software, Formal analysis, Writing–Original Draft, and Writing–Review & Editing.
Jivan Lamichhane: Conceptualization, Data Curation, and Writing–Review & Editing. Abdulaziz Ahmed: Resources, Methodology, Writing–Review & Editing, and Validation. Johnson Thomas: Resources and Writing–Review & Editing. Dursun Delen: Conceptualization, Supervision, and Writing–Review & Editing. William Paiva: Conceptualization, Supervision, Funding Acquisition, and Writing–Review & Editing.
Conflicts of Interest
The authors have no competing interests to declare.
Data and Code Availability
All annotated notes and source code supporting this study have been made publicly
available at https://github.com/hieutrann/ROS_entities_extraction
References

  1. Chung AE, Basch EM. Incorporating the patient’s voice into electronic health records
    through patient-reported outcomes as the “review of systems.” Journal of the American
    Medical Informatics Association. 2015;22(4):914-916. doi:10.1093/jamia/ocu007
  2. Ernecoff NC, Arnold J, Krishnamurti T, et al. Perceptions of Information Transferred in
    Review of Systems Forms: A Qualitative Description. J GEN INTERN MED. Published
    online February 20, 2025. doi:10.1007/s11606-025-09443-4
  3. Jenkins S. History taking, assessment and documentation for paramedics. Journal of
    Paramedic Practice. 2013;5(6):310-316. doi:10.12968/jpar.2013.5.6.310
  4. Phillips A, Frank A, Loftin C, Shepherd S. A Detailed Review of Systems: An Educational
    Feature. The Journal for Nurse Practitioners. 2017;13(10):681-686.
    doi:10.1016/j.nurpra.2017.08.012
  5. Tuite PJ, Krawczewski K. Parkinsonism: A Review-of-Systems Approach to Diagnosis.
    Seminars in Neurology. 2007;27:113-122. doi:10.1055/s-2007-971174
  6. Asadi-Pooya AA, Rabiei AH, Tinker J, Tracy J. Review of systems questionnaire helps
    differentiate psychogenic nonepileptic seizures from epilepsy. Journal of Clinical
    Neuroscience. 2016;34:105-107. doi:10.1016/j.jocn.2016.05.037
  7. Okland TS, Gonzalez JR, Ferber AT, Mann SE. Association Between Patient Review of
    Systems Score and Somatization. JAMA Otolaryngology–Head & Neck Surgery.
    2017;143(9):870-875. doi:10.1001/jamaoto.2017.0671
  8. Everdell E, Borok J, Deutsch A, et al. ATLAS: A positive, high-yield review of patient
    symptoms most significantly associated with melanoma recurrence. Journal of the
    American Academy of Dermatology. 2024;91(6):1118-1124.
    doi:10.1016/j.jaad.2024.07.1516
  9. Podder V, Lew V, Ghassemzadeh S. SOAP Notes. In: StatPearls. StatPearls Publishing;
  10. Accessed April 6, 2025. http://www.ncbi.nlm.nih.gov/books/NBK482263/
  11. Hagan S, Hagan AF. The Review of Systems and the Physical Exam. In: Wong CJ, Jackson
    SL, eds. The Patient-Centered Approach to Medical Note-Writing. Springer International
    Publishing; 2023:153-162. doi:10.1007/978-3-031-43633-8_12
  12. Sinsky C, Colligan L, Li L, et al. Allocation of Physician Time in Ambulatory Practice: A
    Time and Motion Study in 4 Specialties. Ann Intern Med. 2016;165(11):753-760.
    doi:10.7326/M16-0961
  13. Arndt BG, Beasley JW, Watkinson MD, et al. Tethered to the EHR: Primary Care
    Physician Workload Assessment Using EHR Event Log Data and Time-Motion
    Observations. Ann Fam Med. 2017;15(5):419-426. doi:10.1370/afm.2121
  14. Tai-Seale M, Olson CW, Li J, et al. Electronic Health Record Logs Indicate That
    Physicians Split Time Evenly Between Seeing Patients And Desktop Medicine. Health
    Affairs. 2017;36(4):655-662. doi:10.1377/hlthaff.2016.0811
  15. Woolhandler S, Himmelstein DU. Administrative Work Consumes One-Sixth of U.S.
    Physicians’ Working Hours and Lowers their Career Satisfaction. Int J Health Serv.
    2014;44(4):635-642. doi:10.2190/HS.44.4.a
  16. Centers for Medicare & Medicaid Services. Evaluation and Management Services Guide.
    Accessed April 7, 2025. https://www.cms.gov/outreach-and-education/medicare-learningnetwork-
    mln/mlnproducts/mln-publications-items/cms1243514
  17. CPT® Evaluation and Management. American Medical Association. December 27, 2023.
    Accessed April 7, 2025. https://www.ama-assn.org/practice-management/cpt/cptevaluation-
    and-management
  18. Maisel N, Thombley R, Sinsky CA, et al. Primary Care Physician Perceptions of the
    Impact of CMS E/M Coding Changes and Associations with Changes in EHR Time. J GEN
    INTERN MED. Published online February 18, 2025. doi:10.1007/s11606-025-09400-1
  19. Nadkarni PM, Ohno-Machado L, Chapman WW. Natural language processing: an
    introduction. Journal of the American Medical Informatics Association. 2011;18(5):544-
  20. doi:10.1136/amiajnl-2011-000464
  21. Locke S, Bashall A, Al-Adely S, Moore J, Wilson A, Kitchen GB. Natural language
    processing in medicine: A review. Trends in Anaesthesia and Critical Care. 2021;38:4-9.
    doi:10.1016/j.tacc.2021.02.007
  22. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large
    language models in medicine. Nat Med. 2023;29(8):1930-1940. doi:10.1038/s41591-023-
    02448-8
  23. Hu Y, Zuo X, Zhou Y, et al. Information Extraction from Clinical Notes: Are We Ready to
    Switch to Large Language Models? Published online January 7, 2025.
    doi:10.48550/arXiv.2411.10020
  24. Shool S, Adimi S, Saboori Amleshi R, Bitaraf E, Golpira R, Tara M. A systematic review
    of large language model (LLM) evaluations in clinical medicine. BMC Med Inform Decis
    Mak. 2025;25(1):117. doi:10.1186/s12911-025-02954-4
  25. Krishna K, Pavel A, Schloss B, Bigham JP, Lipton ZC. Extracting Structured Data from
    Physician-Patient Conversations by Predicting Noteworthy Utterances. In: Shaban-Nejad
    A, Michalowski M, Buckeridge DL, eds. Explainable AI in Healthcare and Medicine:
    Building a Culture of Transparency and Accountability. Springer International Publishing;
    2021:155-169. doi:10.1007/978-3-030-53352-6_14
  26. Wang Y, Wang L, Rastegar-Mojarad M, et al. Clinical information extraction applications:
    A literature review. J Biomed Inform. 2018;77:34-49. doi:10.1016/j.jbi.2017.11.011
  27. Centers for Medicare & Medicaid Services (CMS). Evaluation & Management Visits.
    Accessed May 6, 2025. https://www.cms.gov/medicare/payment/feeschedules/
    physician/evaluation-management-visits
  28. Denny JC, Miller RA, Johnson KB, Spickard A. Development and Evaluation of a Clinical
    Note Section Header Terminology. AMIA Annu Symp Proc. 2008;2008:156-160.
  29. Denny JC, Spickard A, Johnson KB, Peterson NB, Peterson JF, Miller RA. Evaluation of a
    Method to Identify and Categorize Section Headers in Clinical Documents. J Am Med
    Inform Assoc. 2009;16(6):806-815. doi:10.1197/jamia.M3037
  30. Brown T, Mann B, Ryder N, et al. Language Models are Few-Shot Learners. In: Advances
    in Neural Information Processing Systems. Vol 33. Curran Associates, Inc.; 2020:1877-
  31. Accessed April 12, 2025.
    https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-
    Abstract.html
  32. Hu Y, Chen Q, Du J, et al. Improving large language models for clinical named entity
    recognition via prompt engineering. Journal of the American Medical Informatics
    Association. 2024;31(9):1812-1820. doi:10.1093/jamia/ocad259
  33. Friedl J. Mastering Regular Expressions. O’Reilly Media, Inc.; 2006. Accessed April 16,
  34. https://books.google.com/books?hl=en&lr=&id=GX3w_18-
    JegC&oi=fnd&pg=PR7&dq=regular+expression&ots=PMoiUmdvS-
    &sig=VlE9XrlUzBUyAcGwdDnyyI5boA4
  35. Mistral Small 3.1 | Mistral AI. Accessed April 16, 2025. https://mistral.ai/news/mistralsmall-
    3-1
  36. Azaiz I, Kiesler N, Strickroth S, Zhang A. Open, Small, Rigmarole — Evaluating Llama 3.2
    3B’s Feedback for Programming Exercises. Published online April 1, 2025.
    doi:10.48550/arXiv.2504.01054
  37. Team G, Kamath A, Ferret J, et al. Gemma 3 Technical Report. Published online March 25,
  38. doi:10.48550/arXiv.2503.19786
  39. Topsakal O, Akinci TC. Creating Large Language Model Applications Utilizing
    LangChain: A Primer on Developing LLM Apps Fast. ICAENS. 2023;1(1):1050-1056.
    doi:10.59287/icaens.1127
  40. Chinchor N, Sundheim B. MUC-5 Evaluation Metrics. In: Fifth Message Understanding
    Conference (MUC-5): Proceedings of a Conference Held in Baltimore, Maryland, August
    25-27, 1993. ; 1993. Accessed June 5, 2024. https://aclanthology.org/M93-1007
  41. White C, Dooley S, Roberts M, et al. LiveBench: A Challenging, Contamination-Limited
    LLM Benchmark. Published online April 18, 2025. doi:10.48550/arXiv.2406.19314
  42. Huang L, Yu W, Ma W, et al. A Survey on Hallucination in Large Language Models:
    Principles, Taxonomy, Challenges, and Open Questions. ACM Trans Inf Syst.
    2025;43(2):42:1-42:55. doi:10.1145/3703155
    Supplementary Material: System Prompts and Ollama Configura􀆟on Parameters for
    the Proposed LLM Pipeline
    ROS En􀆟ty Recogni􀆟on:
    FROM … # Specify the model here
    PARAMETER temperature 1
    PARAMETER seed 42
    PARAMETER top_k 10
    PARAMETER top_p 0.5
    SYSTEM “””You are a specialized medical documenta􀆟on AI. You will receive a clinical note. Your
    task is to extract all diseases, symptoms, and body systems (exact original text) and determine
    their posi􀆟ve or nega􀆟ve status based on the context.
    Format your response exactly as shown in the “Output Example” below. If requested to output in
    JSON format, follow the JSON structure given in the “JSON Output Example” below precisely.
    Input Example:
    Mild fever, denies headache, no back pain, GI is nega􀆟ve
    Output Example:
  43. “fever” – posi􀆟ve
  44. “headache” – nega􀆟ve
  45. “back pain” – nega􀆟ve
  46. “GI” – nega􀆟ve
    JSON Output Example:
    [
    {
    “extract”: “fever”,
    “status”: “posi􀆟ve”
    },
    {
    “extract”: “headache”,
    “status”: “nega􀆟ve”
    },
    {
    “extract”: “back pain”,
    “status”: “nega􀆟ve”
    },
    {
    “extract”: “GI”,
    “status”: “nega􀆟ve”
    }
    ]
    Ensure your response strictly follows these formats without devia􀆟on.
    “””
    Body System Classifica􀆟on:
    FROM … # Specify the model here
    PARAMETER temperature 1
    PARAMETER seed 42
    PARAMETER top_k 10
    PARAMETER top_p 0.5
    SYSTEM “””
    You are a specialized medical documenta􀆟on AI that classifies diseases based on their
    associated review of systems (ROS). Your task is to determine the appropriate ROS category for
    a given disease.
    Review of Systems Categories:
    Cons􀆟tu􀆟onal Symptoms
    Eyes
    ENT (Ears, Nose, Throat)
    Cardiovascular
    Respiratory
    Gastrointes􀆟nal
    Genitourinary
    Musculoskeletal
    Integumentary/Breast
    Neurological
    Psychiatric
    Endocrine
    Hematologic/Lympha􀆟c
    Allergic/Immunologic
    Output Format:
    Each disease must be mapped to the most relevant ROS category.
    Format: –>
    Examples:
    Input: “prostate” disease
    Output: prostate –> Genitourinary
    Input: “nausea”
    Output: nausea –> Gastrointes􀆟nal
    Input: “vomi􀆟ng”
    Output: vomi􀆟ng –> Gastrointes􀆟nal
    Input: “diabetes”
    Output: diabetes –> Endocrine
    If the input is not a disease, symptom, body loca􀆟on, or body system, output “None”
    Example:
    Input: “Otherwise”
    Output: None

Lascia un commento

Il tuo indirizzo email non sarà pubblicato. I campi obbligatori sono contrassegnati *