Fine-Tuning Large Language Models
Fine-tuning is the further training of a pretrained language model on task- or domain-specific data so that its general capabilities are specialized for a particular purpose. It is the second stage of the transfer learning paradigm, which pairs a self-supervised pretraining phase on broad data with one or more domain adaptation and fine-tuning steps on task-specific data [^c1]. The dominant pattern in natural language processing is large-scale pretraining followed by adaptation to particular tasks or domains [^c2], and fine-tuning is the mechanism that makes that adaptation possible for classification, generation, extraction, conversation, and code.
The field spans a wide range of techniques. Full fine-tuning updates every weight but is expensive to train and deploy. Parameter-efficient methods — adapters, prefix and prompt tuning, and especially low-rank adaptation (LoRA) and its quantized variant QLoRA — train only a small fraction of parameters; QLoRA reduces memory enough to fine-tune a 65-billion-parameter model on a single 48GB GPU [^c4]. Instruction tuning teaches a model to follow natural-language instructions across many tasks, while preference-based methods such as reinforcement learning from human feedback and direct preference optimization align model behavior with human judgments [^c3]. These approaches are complementary, and modern assistants are typically built by combining them.
Fine-tuning is one of several ways to improve a system's behavior, and choosing among them depends on the failure mode being addressed: prompting changes how the model is instructed, retrieval-augmented generation supplies external information, fine-tuning modifies repeatable model behavior, and distillation compresses a proven workflow [^c5]. Fine-tuning is best suited to stable tasks with measurable behavioral gaps, and it is data-sensitive: scarcity can cause overfitting, poor generalization, and suboptimal performance [^c6].
Practically, the field also encompasses the data used for training, the mechanics of running jobs — hyperparameters, memory optimization, and distributed training — the evaluation of results, the domains where fine-tuning is applied, the tooling that supports it, and the challenges it raises, from compute cost and data privacy to security, labor ethics, and open research questions. This wiki covers each of these facets.