DecodingTrust Framework Wins NeurIPS Award for GPT Model Safety Evaluation
University of Chicago professor Bo Li wins a NeurIPS 2023 award for DecodingTrust, the first comprehensive platform for assessing large language model trustworthiness across eight key perspectives.
University of Chicago professor Bo Li earns the NeurIPS 2023 Outstanding Paper Award for DecodingTrust, a groundbreaking framework that evaluates the trustworthiness of generative pre-trained transformer (GPT) models. As large language models see rapid deployment in sensitive areas like medical diagnoses and financial decisions, this new platform provides the first comprehensive risk assessment to address growing concerns about privacy and safety limitations.
The DecodingTrust framework establishes eight distinct perspectives for evaluating these models: toxicity, stereotype bias, adversarial robustness, out-of-distribution robustness, robustness on adversarial demonstrations, privacy, machine ethics, and fairness. Li states that this organized approach serves as a crucial checklist to help researchers and practitioners comprehensively understand an AI model's abilities and potential vulnerabilities before real-world application.
Testing reveals that while GPT-4 generally outperforms GPT-3.5 across standard trustworthiness metrics, this advantage quickly disappears when users apply misleading prompts or jailbreaking techniques. Because GPT-4 possesses superior instruction-following capabilities, it is actually more susceptible to following malicious or deceptive instructions, highlighting a critical security gap in the most advanced language models available today.