Abstract

The language barrier creates significant healthcare challenges for patients who struggle with English. While large language models (LLMs) have advanced machine translation, their high computational costs, environmental impact, and risk of data breaches have raised concerns. Small language models (SLMs) can address these concerns and be further adapted to specific tasks through parameter-efficient fine-tuning (PEFT) methods, including Quantized Low-Rank Adaptation (QLoRA). This study compares the Spanish-to-English translation quality of two 7-billion-parameter SLMs, Mistral and Llama 2 7B, in the medical domain, using BLEU, chrF, and COMET as evaluation metrics. It further investigates whether QLoRA fine-tuning and Retrieval-Augmented Generation (RAG) on the biomedical domain of the ALIA dataset improve their performance. Results indicate that both models exhibited low lexical accuracy (BLEU, chrF) but high semantic accuracy (COMET). Notably, QLoRA fine-tuning had no impact on lexical and semantic accuracy with the Mistral model, while the fine-tuned Llama 2 model experienced severe repetition. Furthermore, RAG had varying effects on translation performance: the base models showed signs of overfitting, whereas the QLoRA models exhibited no major changes in performance. This work provides an empirical basis for selecting and adapting SLMs for low-resource medical translation, highlighting both their potential and the pitfalls of naive fine-tuning and retrieval augmentation.

Advisor

Rajeev Bukralia

Committee Member

Flint Million

Date of Degree

2026

Language

english

Document Type

Thesis

Degree

Master of Science (MS)

Program of Study

Data Science

Department

Computer Information Science

College

Science, Engineering and Technology

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.

Included in

Data Science Commons

Share

COinS
 

Rights Statement

In Copyright