Future product direction from Franci: track face points locally (MediaPipe is a candidate), let the conversation model create bounded three.js effects, and composite them into the outgoing camera image so other call participants see them. Examples: a hat that follows the face; a floating pie chart answering a spoken question.
Keep this as a follow-up to the current OpenAI/local-processing completion work. Evaluate local landmark tracking, stable coordinates and occlusion, bounded script access to fresh landmarks, latency on the capture path, and loss-of-face cleanup. Reuse the existing transparent renderer and host compositor. No implementation or system-component activation is authorized by this planning item alone.
Acceptance should eventually include another call participant seeing the generated effect. Virtual-camera acceptance remains deferred in the current work.
Future product direction from Franci: track face points locally (MediaPipe is a candidate), let the conversation model create bounded three.js effects, and composite them into the outgoing camera image so other call participants see them. Examples: a hat that follows the face; a floating pie chart answering a spoken question.
Keep this as a follow-up to the current OpenAI/local-processing completion work. Evaluate local landmark tracking, stable coordinates and occlusion, bounded script access to fresh landmarks, latency on the capture path, and loss-of-face cleanup. Reuse the existing transparent renderer and host compositor. No implementation or system-component activation is authorized by this planning item alone.
Acceptance should eventually include another call participant seeing the generated effect. Virtual-camera acceptance remains deferred in the current work.