A Real time Face Emotion Detection System Based on YOLO11 | IJET Volume 12 – Issue 4 | IJET-V12I4P13

IJET
International Journal of Engineering and Techniques
ISSN 2395-1303 · Peer-Reviewed · Open Access
📚 Volume 12, Issue 4
📅 August 2, 2026
📄 Pages 107–113
🔖 ID: IJET-V12I4P13

A Real time Face Emotion Detection System Based on YOLO11

Author(s)

Miss.Bhivsane.P.P, Mr.S.G.Shah

Abstract

Modern Human-Computer Interaction (HCI) trends demand a shift toward emotionally intelligent systems capable of recognizing user affect. Standard computer vision setups typically process environment tasks sequentially, remaining blind to human psychological and physiological variations. This research presents an optimized, end-to-end computer vision engine that unifies a single-stage regression model (YOLO11), a custom multi-layered Convolutional Neural Network (CNN), and dense topological point cloud extractions (MediaPipe FaceMesh) into an asynchronous behavioral analytics framework. The core structural asset of this design centers on a decoupled dual-stream execution loop: ● Stream A utilizes a pruned YOLO11 object boundary network optimized for low-latency face tracking under unconstrained ambient lighting conditions. Once isolated, the sub matrix is normalized and forwarded to a specialized deep CNN to sort the user’s expression across seven key human categories: angry, disgust, fear, happy, neutral, sad, and surprise. ● Stream B executes simultaneously on independent hardware threads, establishing a geometric landmark coordinate array consisting of 468 distinct points. By evaluating spatial vectors across this 3D topographical mesh, the system extracts precise physiological indicators, including involuntary eye closures via the Eye Aspect Ratio (EAR), directional gaze trajectories (LEFT, RIGHT, STRAIGHT), mandibular separation tracking for yawn validation, and relative head orientation changes (UP, DOWN, CENTER). These tracking parameters are aggregated by a multi-parametric fusion engine to calculate a continuous attention score from 0% to 100%. Empirical validation shows that this integrated engine achieves a face-bounding accuracy of 95.1%, an expression classification score of 92.4%, and an EAR blink tracking consistency rate of 93.2%. Natively hosted on commercial edge hardware, the system maintains a stable processing speed of 24 Frames Per Second (FPS) with an end-to-end inference latency bounded tightly at 41 milliseconds. This lightweight footprint removes any reliance on external cloud APIs, ensuring data privacy and making the framework highly viable for driver drowsiness detection, smart learning management systems, and automated medical monitoring.

Keywords

YOLO11, Emotion Recognition, Computer Vision, CNN, Deep Learning, Media Pipe, Face Mesh, Human Behavior Analysis, Artificial Intelligence, Real-Time Detection, Lightweight Regression Networks, Affective Computing, Geometrical Coordinate Mapping, Human-Computer Interaction, Visual Telemetry, Attention Estimation.

Conclusion

This project successfully designed, validated, and deployed a real-time multi-parametric facial emotion and behavioral analysis pipeline optimized for edge device execution. By combining the single-stage localization of YOLO11 with a custom CNN and Media Pipe Face Mesh sub-routines, the architecture achieved a stable processing speed of 24 FPS with an exceptionally low latency of 41 milliseconds. Quantitative validations confirmed a 95.1% face bounding accuracy and a 92.4% expression classification score, meeting the strict requirements for safety-critical deployment. The project provides three key contributions to the domain of human-computer interaction: A Decoupled Asynchronous Multi-Threading Framework that isolates GPU-bound neural processing from CPU-bound geometric extractions to eliminate performance bottlenecks. A Multi-Metric Analytics Fusion Engine that resolves single-variable context blindness by cross-referencing emotional labels with scale-invariant eye and head vectors. An Edge-Deployed Architecture that requires less than 500 MB of VRAM, preserving user data privacy and removing cloud API dependencies. The resulting software engine successfully bridges psychological frameworks with advanced computer vision, establishing a reliable foundation for emotionally intelligent computing systems. FUTURE WORK: Emotion Forecasting: Predict the probable next emotional state based on recent emotional transitions rather than only recognizing the current emotion. Real-time alert systems are designed to monitor environments continuously and deliver immediate notifications when anomalies, threats, or emergencies are detected. The ultimate evolution of affective computing lies in Multi-Modal Data Fusion. Future iterations of this architecture willing corporate real-time acoustic processing Streams to complement the visual streams. Transformer-based deep learning model

References

[1]R.W.Picard, Affective Computing. Cambridge, MA, USA: MIT Press, 1997.
[2]P. Ekman and W. V. Friesen, Facial Action Coding System: A Technique for the Measurement of Facial Movement. Palo Alto, CA, USA: Consulting Psychologists Press, 1978.
[3]J.Redmon,S.Divvala,R.Girshick,andA.Farhadi,"YouOnlyLookOnce:Unified, Real-Time Object Detection," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 779-788.
[4]Ultralytics, "YOLO11 Architecture and Real-Time Object Detection Engine Optimization Documentation, "Ultralytics Deep Learning Repository, 2024.[Online]. Available: https://github.com/ultralytics/ultralytics
[5]P. Viola and M. Jones, "Rapid Object Detection using Boosted Cascade of Simple Features, “in Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), Kauai, HI, USA, 2001, pp. I-511–I-518.
[6]M.Abadietal.”Tensor Flow:e-Scale Machine Learning on Heterogeneous Distributed Systems," arXiv preprint arXiv:1603.04467, 2016.
[7]D. P. Kingman and J. Ba, "Adam: A Method for Stochastic Optimization," in Proceedings of the 3rd International Conference on Learning Representations (ICLR), San Diego, CA, USA, 2015.
[8]C.Lugaresietal.”MediaPipe: A Framework for Building Perception Pipelines, "arXiv preprint arXiv: 1906.08172, 2019.
[9]T. Soukupova and J. Cech, "Real-Time Eye Blink Detection Using Facial Landmarks, “in Proceedings of the 21st Computer Vision Winter Workshop (CVWW), Rimske Toplice, Slovenia, 2016, pp. 1-8.
[10]G.Bradski, "The OpenCV Library,"Dr. Dobb’s Journal of Software Tools, vol.25, Pp.120-123,2000.
[11]I. J. Good fellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016.
[12]Z.Zheng,P.Wang,W.Liu,J.Li,R.Ye,andD.Ren,"Distance-IoULoss:Faster and Better Learning for Bounding Box Regression," in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, pp. 12565-12572, 2020.
[13]J. Redmon et al., “YOLO: Real-Time Object Detection,” IEEE Conference on Computer Vision and Pattern Recognition.
[14]F. Larradet, R. Niewiadomski, G. Barresi, D. G. Caldwell, and L. S. Mattos, ‗‗Toward emotion recognition from physiological signals in the wild: Approaching the methodological issues in real-life data collection, Frontiers Psychol., vol. 11, p. 1111, Jul. 2020.
[15]Dahl R.E., Harvey A.G. Sleep in children and adolescents with behavioral and emotional disorders https://www.sciencedirect.com/science/article/pii/S1556407X07000513
[16]Wilson G.F., Russell C.A. Real-time assessment of mental workload using psycho physiological measures and artificial neural networks Hum. Factors, 45 (4) (2003), pp. 635-644 Google Scholar
[17]Kang S., Park C.Y., Kim A., Cha N., Lee U.Understanding emotion changes in mobile experience CHI ’22, Association for Computing Machinery, New York, NY, USA (2022), 10.1145/3491102.3501944

[18]V.Lepetit,F.Moreno-Noguer,andP.Fua,"EPnP:AnAccurateO(n)Solutiontothe PnP Problem, “International Journal of Computer Vision, vol. 81, no. 2, pp. 155-166, 2009
[19]A Real Time Face Emotion Detection System Based On YOLO11 | IJET
[20]V.Lepetit,F.Moreno-Noguer,andP.Fua,"EPnP:AnAccurateO(n)Solutiontothe PnP Problem, “International Journal of Computer Vision, vol. 81, no. 2, pp. 155-166, 2009
[21]R. Manzari et al., "BEFUnet: Hybrid CNN-Transformer Architecture for Medical Image Segmentation," arXiv, 2024.
[22]Y. Sun et al., "DA-TransUNet: Dual Attention Transformer Network," arXiv, 2023.
[23]J. Zhang et al., "Medical Image Segmentation: A Review," IEEE Access, 2025.
[24]S. He et al., "Deep Learning in Medical Image Classification," Frontiers in Medicine, 2025.
[25]A. Singh et al., "Medical Image Enhancement Techniques," Springer, 2024.
[26]The Authors. This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/
[27]F. Larradet, R. Niewiadomski, G. Barresi, D. G. Caldwell, and L. S. Mattos, Toward emotion recognition from physiological signals in the wild: Approaching the methodological issues in real-life data collection,‘‘ Frontiers Psychol., vol. 11, p. 1111,

📋 How to Cite This Paper

Miss.Bhivsane.P.P, Mr.S.G.Shah (2026). A Real time Face Emotion Detection System Based on YOLO11. International Journal of Engineering and Techniques, 12(4), 107–113. ISSN: 2395-1303. DOI: https://doi.org/10.5281/zenodo.21760936
© 2026 International Journal of Engineering and Techniques (IJET). All rights reserved. · ijetjournal.org
Submit Your Paper