AI Translates Preference Models into Natural Language for Editing.

Zachary Wojtowicz, Ayush Nayak, Jacob Andreas· July 21, 2026 View original

Summary

This method, "weights to words," automatically discovers and describes preference dimensions in natural language from choice data, pairing them with model vectors. It allows users to inspect and edit AI preference model inferences in real-time, improving prediction accuracy and user trust across diverse domains like moral dilemmas and movie selection.

Statistical learning algorithms are increasingly used to infer human preferences from complex choice data. However, the opacity of these models makes it difficult to understand which factors truly drive decisions and to correct errors. A new method, termed "weights to words," aims to bridge this gap by making AI preference models more interpretable and editable. This method takes a dataset of choice problems and automatically identifies relevant preference dimensions within the domain. Each dimension is then described in natural language and linked to a corresponding vector in the model's internal representation space. This dual representation addresses both the under-determination of factors influencing choices and the opacity of the models. Users can inspect the model's inferences in natural language and make real-time edits to these preference profiles. Experiments across various domains, including moral dilemmas and movie selection, show significant benefits. Regularizing a preference model towards the learned basis increases prediction accuracy on unseen choices, and incorporating user-structured edits further enhances this accuracy. Participants in human-subjects experiments preferred the method's inferred preference profiles and found its predictions more accurate, highlighting its potential for improving trust and utility in AI-driven preference systems.

Why it matters

For product managers, marketers, and AI developers, this innovation offers a way to build more transparent, controllable, and accurate preference models, leading to better personalized recommendations, improved user experience, and increased trust in AI systems.

How to implement this in your domain

  1. 1Integrate "weights to words" into your recommendation engines or personalization platforms to enhance transparency.
  2. 2Develop user interfaces that allow customers or internal teams to inspect and edit preference dimensions in natural language.
  3. 3Apply this method to improve the accuracy of preference models by incorporating user feedback and structured edits.
  4. 4Utilize the natural language descriptions to better understand and debug AI-driven decision-making processes.
  5. 5Explore new product features based on explicit, editable preference profiles for enhanced user control.

Who benefits

E-commerceMedia & EntertainmentMarketingProduct DevelopmentCustomer Service

Key takeaways

  • AI preference models are often opaque, making it hard to understand and correct their inferences.
  • "Weights to words" translates model preferences into natural language dimensions.
  • Users can inspect and edit these natural language preference profiles in real-time.
  • This method improves prediction accuracy and user trust in AI-driven personalization.

Original post by Zachary Wojtowicz, Ayush Nayak, Jacob Andreas

"arXiv:2607.16232v1 Announce Type: new Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is gene…"

View on X

Originally posted by Zachary Wojtowicz, Ayush Nayak, Jacob Andreas on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses