International Journal of Engineering
Trends and Technology

Research Article | Open Access | Download PDF
Volume 74 | Issue 9 | Year 2026 | Article Id. IJETT-V74I9P122 | DOI : https://doi.org/10.14445/22315381/IJETT-V74I9P122

Facial Expression-based Music Recommendation using Fourier Image Transformer


Raghunathan S, Swati Sharma

Received Revised Accepted Published
17 Feb 2026 25 Jul 2026 07 Aug 2026 30 Sep 2026

Citation :

Raghunathan S, Swati Sharma, "Facial Expression-based Music Recommendation using Fourier Image Transformer," International Journal of Engineering Trends and Technology (IJETT), vol. 74, no. 9, pp. 304-332, 2026. Crossref, https://doi.org/10.14445/22315381/IJETT-V74I9P122

Abstract

The ability to recognize human facial expressions is a fundamental requirement for developing advanced human–computer interaction systems. Humans convey emotions primarily through facial expressions, which has driven the development of Facial Expression Recognition (FER) technology in areas such as affective computing and personalized recommendation systems. By leveraging FER, systems can continuously monitor changes in a user’s emotional state throughout the day and generate context-aware responses, thereby enhancing overall user experience. The Spotify application of FER is a music recommendation, where songs are suggested based on facial expressions to match or influence the user’s emotions. Traditional FER approaches primarily rely on spatial-domain features, which often struggle to capture global context and subtle frequency-based patterns in facial images. To address these limitations, this study introduces the Fourier Image Transformer (FIT), a frequency-domain transformer model capable of learning discriminative global representations of facial expressions. By transforming facial images into the frequency domain, FIT captures both local and global variations in facial patterns, improving recognition accuracy and robustness. The proposed framework is trained and tested on the Real-world Affective Faces Multi-Label (RAF-ML) dataset and FER2013, which contains diverse real-world facial expressions under varying conditions. The preprocessing steps, including face detection, image enhancement, resizing, and normalization, to ensure reliable feature extraction. Recognized FER are subsequently mapped to selected music tracks, enabling personalized and image-based music recommendation. This approach offers a seamless, non-intrusive interface that adapts to user emotions in real-time. The FIT model achieved outstanding performance with an accuracy of 98.97%, precision of 98.92%, recall of 98.87%, and an F1-score of 98.90%. The study demonstrates how frequency-domain transformers can enhance FER performance and support intelligent, emotion-driven applications.

Keywords

Fourier Image Transformer, Real-world Affective Faces Multi-Label) dataset, FER2013, Image enhancement, Resizing, and Normalization.

References

[1] Thomas Kopalidis et al., “Advances in Facial Expressions Recognitions: A Survey of Methods, Benchmark, Models, and Dataset,” Information, vol. 15, no. 3, pp. 1-61, 2024.
[CrossRef] [Google Scholar] [Publisher Link]             

[2] Ruchi Jayaswal et al., “Advances in Facial Expressions Recognitions Technologies for Emotions Analysis,” Discover Computing, vol. 28, no. 1, pp. 1-56, 2025.
[CrossRef] [Google Scholar] [Publisher Link]   

[3] Pit Pichappan, “A Review of the Emotions-Induced Music Recommendations System,” Journal of Digital Information Management, vol. 23, no. 2, pp. 112-133, 2025.
[CrossRef] [Publisher Link]    

[4] Sheetal Patil et al., “Review on Music Emotions Analysis using Machine Learnings: Technologies, Method, Dataset, and Challenge,” Discover Applied Sciences, vol. 7, no. 7, pp. 1-14, 2025.
[CrossRef] [Google Scholar] [Publisher Link]    

[5] Vilas Gaikwad et al., “Emotions-Driven Music Recommendations Systems,” SAMRIDDHI: A Journal of Physical Science, Engineering and Technology, vol. 16, no. 4, pp. 136-140, 2024.
[CrossRef] [Publisher Link] 

[6] Rashini Liyanarachchi, Aditya Joshi, and Erik Meijering, “A Survey on Multimodal Music Emotions Recognitions,” arXiv, pp. 1-26, 2025.
[CrossRef] [Google Scholar] [Publisher Link]   

[7 ]Sepideh Kalateh et al., “A Systematic Review on Multimodal Emotions Recognitions: Building Block, Current states, Application, and Challenge,” IEEE Access, vol. 12, pp. 103976-4019, 2024.
[CrossRef] [Google Scholar] [Publisher Link]    

[8] Rosa A. García-Hernández et al., “A Systematic Literature Review of Modalities, Trend, and Limitation in Emotions Recognitions, Affective Computing, and Sentiments Analysis,” Applied Sciences, vol 14, no. 16, pp. 1-25, 2024.
[CrossRef] [Google Scholar] [Publisher Link]    

[9] You Wu, Qingwei Mi, and Tianhan Gao, “A Comprehensive Review of Multimodal Emotions Recognitions: Technique, Challenges, and Future Direction,” Biomimetics, vol 10, no. 7, pp. 1-24, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[10] Manju Priya Arthanarisamy Ramaswamy, and Suja Palaniswamy, “Multimodal Emotions Recognitions: A Comprehensive Review, Trend, and Challenges,” Wiley Interdisciplinary Review: Data Mining and Knowledge Discovery, vol. 14, no. 6, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[11] Guimin Hu et al., “Recent Trend of Multimodal Affective Computing: A Survey from NLP Perspectives,” arXiv, pp. 1-22, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[12] Erkang Jing et al., “Emotions-Aware Personalized Music Recommendations with a Heterogeneity-Aware Deep Bayesians Network,” ACM Transaction on Information System, vol. 43, no. 5, pp. 1-43, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[13] Madipally Sai Krishna Sashank et al., “Mood-based Music Recommendations Systems using Facial Expressions Recognitions and Texts Sentiments Analysis,” Journal of Theoretical and Applied Information Technology, vol. 100, no. 19, 2022. [Publisher Link]

[14] Rajesh B et al., “Music Recommendations based on Facial Emotions Recognitions,” arXiv, pp. 1-8, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[15] Hung Nguyen et al., “A Model for Songs Recommendations based on Facial Emotions Analysis and Musical Emotions,” International Journal of Intelligent Engineering and System, vol. 17, no. 4, pp. 1028-1041, 2024.
[CrossRef] [Google Scholar]

[16] Fuyan Ma, Bin Sun, and Shutao Li, “Facial Expressions Recognitions with Visual Transformer and Attentional Selective Fusions,” IEEE Transaction on Affective Computing, vol. 14, no. 2, pp. 1236-1248, 2021.
[CrossRef] [Google Scholar] [Publisher Link]

[17] Mei Bie et al., “Swin-FER: Swin Transformers for Facial Expressions Recognitions,” Applied Sciences, vol. 14, no. 14, pp. 1-14, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[18] Janhavi Kaimal et al., “Real Time Emotions Based Music Players,” International Journal for Research in Applied Science and Engineering Technology (IJRASET), vol. 12, no. 4, pp. 5601-5607, 2024.
[CrossRef] [Publisher Link]

[19] V.S.G.S. Phaneendra Bottu, and K. Ragavan, “Emotions-based Music Recommendations Systems Integrating Facial Expressions Recognitions and Lyrics Sentiment Analysis,” IEEE Access, vol. 13, pp. 87740-87752, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[20] Nilesh Bhaskarrao Bahadure et al., “Machine Learning-Based Music Classifications and Recommendations Systems from Spotify,” International Journal of Computer Information Systems and Industrial Management Applications, vol. 17, pp. 143-154, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[21] Nandini Gupta et al., “Intelligent Music Recommendations Systems based on Face Emotions Recognitions,” 2023 International Conferences on Computing, Communications, and Intelligent System (ICCCIS), Greater Noida, India, pp. 110-115, 2023.
[CrossRef] [Google Scholar] [Publisher Link]

[22] Shalaka Prasad Deore, “SongRec: A Facial Expression Recognitions System for Songs Recommendations using CNN,” International Journal of Performability Engineering, vol. 19, no. 2, pp. 115-121, 2023.
[CrossRef] [Google Scholar] [Publisher Link]

[23] Mst. Lima Akter Asha et al., “Suggesting Playlists and Playing Preferred Music based on Emotions from Facial Expressions,” 2024 3rd International Conferences for Innovations in Technology (INOCON), Bangalore, India, pp. 1-5, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[24] Tanisha Kapoor, Arnaja Ganguly, and D Rajeswari, “Music Recommendations based on Facial Expression using Data Augmentations,” 2023 3rd International Conferences on Innovative Mechanism for Industry Application (ICIMIA), Bengaluru, India, pp. 1172-1178, 2023.
[CrossRef] [Google Scholar] [Publisher Link]

[25] Dhruv Chopra et al., “Feel-Tunes: An Emotion-Driven Music Recommender System,” Proceedings of the International Conference on Innovative Computing & Communication (ICICC 2024), pp. 1-6, 2023.
[
Google Scholar] [Publisher Link]

[26] Ashish Tripathi et al., “Facial Emotions-based Songs Recommender Systems using CNN,” International Journal of Engineering Trends and Technology, vol. 72, no. 6, pp. 315-327, 2024.
[CrossRef] [Publisher Link]

[27] Narayan Paudel et al., “Facial Emotions Recognition System using CNN for Songs Mapping,” International Journal on Engineering Technology, vol. 1, no. 2, pp. 312-323, 2024.
[CrossRef] [Google Scholar] [Publisher Link]

[28] Porawat Visutsak et al., “Mood-based Music Discovery: A System for Generating Personalized Thai Music Playlist using Emotions Analysis,” Applied System Innovation, vol. 8, no. 2, pp. 1-24, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[29] Abdoul Malik, Mohammad Adnan Ayoubi, and Mohammad Alsuedani, “Deep Learning for Facial Emotions Recognitions on the FER-2013 Dataset using the ResNet50v2 Model,” OMU Journal of Engineering Sciences and Technology, vol. 5, no. 2, pp. 19-33, 2025.
[Google Scholar] [Publisher Link]

[30] Baiyi Zhang et al., “Diversecer-Net: A Compound Emotions Recognitions Framework based on Residual Anti-Aliasing Attentions Networks and Diverse Data Augmentations,” Biomedical Signal Processing and Control, vol. 118, pp. 1-16, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[31] Navid Falah, Behnam Yousefimehr, and Mehdi Ghatee, “Predicting Music Tracks Popularity by Convolutional Neural Network on Spotify Feature and Spectrograms of Audio Waveforms,” arXiv, pp. 1-12, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[32] Florian Schwarzhans et al., “Image Normalizations Technique and their Effects on the Robustness and Predictive Powers of Breast MRI Radiomics,” European Journal of Radiology, vol. 187, pp. 1-12, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[33] Rakhmonalieva Farangis Oybek Kizi, Tagne Poupi Theodore Armand, and Hee-Cheol Kim, “A Review of Deep Learning Technique for Leukemia Cancer Classifications based on Blood Smear Image,” Applied Biosciences, vol. 4, no. 1, pp. 1-32, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[34] Muhammad Shahrul Zaim Ahmad et al., “Impact of Images Enhancements using Contrast-Limited Adaptive Histogram Equalizations (CLAHE), Anisotropic Diffusions, and Histogram Equalizations on Spine X-Ray Segmentations with U-Net, Mask R-CNN, and Transfer Learnings,” Algorithms, vol. 18, no. 12, pp. 1-22, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[35] Tim-Oliver Buchholz, and Florian Jug, “Fourier Image Transformer,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Orleans, LA, USA, pp. 1846-1854, 2022.
[CrossRef] [Google Scholar] [Publisher Link]

[36] S. Kumar Reddy Mallidi, and Rajeswara Rao Ramisetty, “Advancements in Training and Deployments Strategies for AI-based Intrusions Detections System in IoT: A Systematic Literature Review,” Discover Internet of Things, vol. 5, no. 8, pp. 1-33, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[37] Xueer Sun et al., “Emotion-Aware Cross-Modal Music Generations based on Multimodal Emotions Recognitions,” Alexandria Engineering Journal, vol. 133, pp. 254-270, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[38] Yaoyu Sun, “Feature Centric based Deep Learning Approach for Music Moods Recognitions with HuBERT Transformer Models,” Scientific Reports, vol. 15, no. 1, pp. 1-34, 2025.
[CrossRef] [Google Scholar] [Publisher Link]

[39] Yassine El Boudouri, and Amine Bohi, “Emonext: An Adapted Convnext for Facial Emotions Recognitions,” 2023 IEEE 25th International Workshops on Multimedia Signals Processing (MMSP), Poitiers, France, pp. 1-6, 2023.
[CrossRef] [Google Scholar] [Publisher Link]

[40] Bettina Shirley Richard et al., “Vivify—Emotion Recognitions and Music Recommendations using Transformer,” IEEE Access, vol. 13, pp. 204894-204907, 2025.
[CrossRef] [Google Scholar] [Publisher Link]