Police narrative reports from 911 responses contain valuable signs for early behavioral health intervention, but obtaining them manually is time-consuming and tedious. We present RECAP, a human-AI collaboration framework that fine-tunes three large language models: Mistral-7B, Llama-3-8B, and TinyLlama to classify BH cases in narratives and generate short, sentence-level rationales. The models are trained on an annotated corpus using supervised instruction data and then aligned using direct preference optimization (DPO), allowing patrol officer preferences to continually influence system behavior without disrupting existing workflows. On a held-out test set, Mistral-7B achieves 85.2% weighted accuracy and 84.2% F1-score, matching strong prior baselines while improving interpretability by short, text-span-linked explanations; Llama-3-8B performs similarly, while TinyLlama provides competitive accuracy at a lower compute cost. RECAP is designed to reduce manual effort and identify behavioral health signs earlier in public narratives, while providing rationales and maintaining officer control.
Stigall, W. A., Nweke, F., Walker, H. N., Khan, M. A. A. H., Perry, S., Pei, Y., Thomas, D., Nandan, M.
Applied Computing and Intelligence, 5(2), pp. 337-347
2025
Abstract