Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach

Amrita Shaw, Chandrasekar S. N., Sai Muthukumar V., Jhinuk Gupta, Deepak L. N. Kallepalli

arXiv:2608.20440 · 2026-08-24 공개 · arXiv · PDF

feature-selection decision-trees spectral-decomposition k-means-clustering food-authentication frugal-ai nnls raman-spectroscopy

Abstract

Authentication of edible oils in processed foods is important for food quality, fraud prevention, and regulatory compliance. This study establishes an integrated Raman spectroscopy and machine-learning framework that links intrinsic spectral organization, interpretable classification, and Physics-Informed Artificial Intelligence (PI-AI). Five edible oils were investigated in pure form and within a fried-potato-chip matrix using t-SNE, K-means clustering, Decision Trees, and Non-Negative Least Squares (NNLS)-based spectral decomposition. Unsupervised analyses revealed substantially stronger class organization and separability in pure oils, whereas food-matrix effects introduced pronounced spectral overlap. Decision Trees achieved 100% classification accuracy for pure oils using only four Raman variables from the original 1866-feature spectral space. These four variables, consistently identified by both pre-pruned and post-pruned models, represented only approximately 0.21% of the available spectral information while retaining perfect test-set performance. For matrix-containing samples, NNLS-based PI-AI spectral decomposition substantially improved classification by separating oil-related signatures from paper and potato contributions. Optimized post-pruned models achieved accuracies of 86.4% and 85.4% for paper-subtracted and paper-plus-potato-subtracted datasets, respectively, while reducing the number of important Raman variables to only five and four. The compact four-feature representation further reduced the data footprint by 99.44% without loss of classification accuracy. Collectively, these findings demonstrate that accurate Raman-based oil identification can be achieved through physically meaningful, highly compact, and interpretable spectral representations, providing a promising foundation for Frugal AI, Edge AI, portable sensing, and embedded food-quality monitoring.

한국어 요약

한 줄 요약

람ان 분광과 결정 트리, K-평균 클러스터링을 결합한 물리 기반 AI를 활용한 식용유 인증 방법을 제시한다.

핵심 기여도

핵심 아이디어

이 연구는 람안 분광 데이터를 해석할 때, 전통적인 머신러닝과 물리 기반 AI를 결합하여 해석 가능성과 정확도를 동시에 확보하려는 접근을 제시한다. 결정 트리는 복잡한 1866차원의 스펙트럼 데이터를 4개의 핵심 변수로 압축하면서도 100% 정확도를 유지함으로써, 해석 가능한 모델 설계의 가능성을 보여준다. 또한, NNLS 기반의 스펙트럼 분해는 종이와 감자 매트릭스의 영향을 제거함으로써, 오일 관련 스펙트럼만 분리해 분류 정확도를 높이는 데 기여한다. 이는 물리적 의미를 갖는 스펙트럼 표현이 머신러닝 성능 향상에 중요한 역할을 할 수 있음을 시사한다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 해석 가능하고 물리적으로 의미 있는 스펙트럼 표현을 통해 AI 모델의 성능과 신뢰성을 동시에 향상시킬 수 있음을 보여준다. 특히, 결정 트리와 NNLS 기반 분해 기법은 Frugal AI 및 Edge AI 적용에 유리한 점이 있다. 그러나 연구는 5종의 오일에만 적용되었으며, 더 다양한 오일 종류나 매트릭스 조건에서의 일반화 가능성은 추가 연구가 필요하다. 또한, 분류 정확도가 감자칩 매트릭스 내에서는 85% 수준으로, 실용적 적용 시 오류를 줄이는 방안이 필요하다.

실용적 활용

이 방법은 식품 품질 모니터링, 식품 사기 방지, 이동식 센서 및 임베디드 시스템에 적용 가능하다. 특히, 해석 가능한 AI 모델은 규제 준수와 신뢰성 확보에 기여할 수 있다.