75%VRAM Reduction
UnslothPEFTQLoRARAGSFT
Fine-tuning Qwen2.5-3B-Instruct
Fine-tuned Qwen2.5-3B-Instruct under tight hardware limits via 4-bit QLoRA (Quantized Low-Rank Adaptation)and paged optimizers to power a memory-efficient CRAG pipeline, optimizing structured verification logic and hallucination-free generation.