Controllable Expressivity for Three-Dimensional Talking Avatars
PDF

Keywords

3D avatars
gaze
multimodal annotation

Abstract

This internal reference article examines controllable expressivity in three-dimensional talking avatars through a design-and-assurance lens. It synthesizes the allocated target literature without reporting new experiments, observations, or performance estimates. The analysis treats the practical unit of review as a speech signal, facial and ocular controls, rendering pipeline, audience judgment, and release gate. That framing keeps technical mechanisms, evidence quality, user consequences, and institutional controls visible in the same argument. Particular attention is given to how modular control can support expression, evaluation, and safety. The review distinguishes what each cited source directly addresses from the cross-domain principles used for internal comparison. It argues that credible adoption depends on traceable requirements, context-sensitive evaluation, explicit uncertainty, and a documented path for human intervention. The result is a structured reference for teams considering telepresence, digital instruction, and accessible interfaces, especially where uncanny behavior, identity misuse, and aesthetic bias could turn a technically plausible component into an unreliable system. The article is intended to support scoping, design review, and evidence planning; it is not a claim of product readiness or an original empirical study.

PDF

References

Cudeiro, D., Bolkart, T., Laidlaw, C., Ranjan, A., & Black, M. J. (2019). Capture, learning, and synthesis of 3D speaking styles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10101–10111). IEEE.

Liu, K., Xiong, H., Zhang, J., & Peng, M. (2026). MOSAIC: A Cognitively Motivated Multi-Agent Framework for Interpretable and Training-Free Empathetic Dialogue. Electronics, 15(10), 2078.

Liu, K., Xiong, H., Zhang, J., & Peng, M. (2026). Unifying Aesthetic Evaluation via Multimodal Annotation and Fine-Grained Sentiment Analysis. Big Data and Cognitive Computing, 10, 37.

Prajwal, K. R., Mukhopadhyay, R., Namboodiri, V. P., & Jawahar, C. V. (2020). A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM International Conference on Multimedia (pp. 484–492). Association for Computing Machinery. https://doi.org/10.1145/3394171.3413532

Richard, A., Zollhöfer, M., Wen, Y., de la Torre, F., & Sheikh, Y. (2021). MeshTalk: 3D face animation from speech using cross-modality disentanglement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 1173–1182). IEEE.

Suwajanakorn, S., Seitz, S. M., & Kemelmacher-Shlizerman, I. (2017). Synthesizing Obama: Learning lip sync from audio. ACM Transactions on Graphics, 36(4), Article 95. https://doi.org/10.1145/3072959.3073640

Tao, J., Lyu, R., & Cao, X. (2026). A Deep Learning-Based Automated Content Moderation Framework for Online Platforms. Future-Adaptive Intelligence and Lifelong Systems, 1(1).

Thies, J., Zollhöfer, M., Stamminger, M., Theobalt, C., & Nießner, M. (2016). Face2Face: Real-time face capture and reenactment of RGB videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 2387–2395). IEEE. https://doi.org/10.1109/CVPR.2016.262

Xiao, F., Song, Y., & Chen, J. (2026, May). DLA-Net: Dynamic Label-Aware Network with Dual-Branch Feature Fusion for Mixed-Type Wafer Map Defect Recognition. In 2026 6th International Conference on Electronics, Circuits and Information Engineering (ECIE) (pp. 72-78). IEEE.

Xiong, H., Zhang, J., Wang, Z., Pan, T., & Hu, Q. (2026). VividTalker: A Modular Framework for Expressive 3D Talking Avatars with Controllable Gaze and Blink. In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Copyright (c) 2026 Daniel Morgan, Thomas Miller (Author)