Abstract
A new model-based coding system for video-telephone and video-conference systems is proposed. In this system, a facial image in the remote terminal of a speaker without any specific expression is transmitted over a telephone line, and a wireframe model of the face is made and stored in the receiving terminal in advance of a conversation. During the conversation, displacement of 34 feature points on the speaker's face is measured and transmitted to the receiving terminal. Also, by modifying the wireframe and model image with this displacement data, the speaker's facial image is reconstructed. Parallel with the displacement data, isodensity lines delineating equal gray level picture elements in the facial image are transmitted and used to compensate density errors in the reconstructed images. There images are generated by changes of surface angles between the facial parts and the illumination source resulting from facial expressions.
This system requires 19 kbps to transmit the facial images, and the quality of the output images in this system is excellent.