07 August 2026
New Publication: Performance of Large Language Models in Classifying Ultra-Processed Foods Evaluated
A study conducted by Assoc. Prof. Hatice Merve Bayram and Prof. Dr. Arda Ozturkcan from the Department of Nutrition and Dietetics at Istanbul Gelisim University, together with Assoc. Prof. Sedat Arslan from the Department of Nutrition and Dietetics at Bursa Uludag University, has been published in the Q1 journal International Journal of Food Science and Technology.
The study, entitled "Can Machines Detect Ultra-Processed Foods? A Head-to-Head Evaluation of Large Language Models Using NOVA Classification," evaluated how accurately large language models (LLMs) can classify ultra-processed foods according to the NOVA classification system.
The research analyzed 2,920 packaged food products sold in Türkiye's three largest supermarket chains, representing 53.2% of the national retail market. First, all products were independently classified according to the NOVA system by two trained dietitians to establish the reference standard. The classification performance of ChatGPT 5.2, Gemini 3, and Grok 4.1 was then compared against this expert reference.
The findings showed that, when using the standard prompt, all three AI models underestimated the prevalence of ultra-processed foods compared with the expert classification. Among the evaluated models, ChatGPT 5.2 demonstrated the best performance, achieving the highest accuracy, specificity, and F1 score. Nevertheless, none of the models produced results sufficiently consistent to replace expert assessment.
The study also investigated the influence of prompt design on model performance. Simpler and more concise prompts significantly improved the classification accuracy of both ChatGPT and Gemini compared with more detailed prompts. Furthermore, repeating the same prompts five months later revealed substantial differences in classification outcomes, highlighting the impact of model updates over time.
The findings suggest that large language models have considerable potential to support researchers in identifying ultra-processed foods. However, reliable application requires carefully designed prompts, regular calibration, and expert validation. The study is expected to contribute to the development of AI-assisted nutrition research and automated food classification systems.
Read the full article:
https://doi.org/10.1093/ijfood/vvag142