top of page

Meta Holograms Put a Synthetic Face on Ray-Ban Display Calls

Sep 27
12 min read

Meta Holograms will enter Early Access this fall, giving US Meta Ray-Ban Display users a synthetic face during WhatsApp video calls. The first release shows only the caller from the shoulders up. More importantly, it generates that appearance from the caller’s voice instead of capturing the person’s face during the call.

That distinction turns an ordinary glasses feature into a test of synthetic presence. The person receiving the call sees a photorealistic digital version of the wearer, complete with inferred expressions. The image looks like video, but it is generated by a model running on Meta’s servers.

Apple established the closest consumer reference with Vision Pro Personas, which represent headset wearers during FaceTime and supported conferencing calls. Meta is taking a different route. It plans to put its first consumer Hologram inside everyday-looking glasses, while using WhatsApp as the initial distribution channel.

The result is more accessible than a spatial avatar inside a headset, but also more ambiguous. Meta must prove that callers want a plausible synthetic face when a normal video feed is unavailable, and that recipients can understand what they are seeing.

Meta Holograms Start With WhatsApp, Not Virtual Reality

Meta is launching its realistic-avatar system as a practical workaround for camera placement, not as a fully spatial telepresence product.

Meta announced Hologram at Connect 2026 as one of several additions to Meta Ray-Ban Display. According to the company’s Ray-Ban Display update, the feature will begin rolling out through Early Access in the United States later this fall.

The initial use case is straightforward. Smart glasses support hands-free calls, but their outward-facing camera cannot provide a conventional view of the wearer. A caller can show another person what lies ahead, yet the recipient cannot see the caller’s face.

Hologram fills that missing camera angle with a generated portrait. The wearer completes a short capture using the Meta AI phone app before making calls. The process records facial appearance and examples of the user speaking expressively.

During a WhatsApp call, the glasses provide an audio stream rather than a live facial view. A generative diffusion model, which creates visual frames by progressively synthesizing image detail, runs on Meta’s servers. It uses the voice to infer facial movements and sends the resulting two-dimensional video to the recipient.

Meta’s public description says expressions are inferred from both what a caller says and how the person says it. That means the model does not observe every movement it portrays. It predicts a visual performance that appears compatible with the incoming speech.

The recipient does not need to wear glasses or enter a virtual environment. The Hologram appears inside a regular WhatsApp video call on the other person’s device. That broad compatibility is central to Meta’s approach.

The first release also carries clear limitations. The capture preserves the clothing and background recorded during setup, according to hands-on reporting. Changing that presentation requires another capture rather than an ordinary wardrobe change or camera adjustment.

Meta is exploring additional inputs for later versions. Data from the glasses’ inertial measurement unit, or IMU, could connect the avatar’s head rotation to the wearer’s real movement. The company is also considering camera-assisted backgrounds and generated clothing changes.

None of those additions has a public release date. The fall version remains a shoulders-up, voice-driven representation with a preselected appearance and environment.

That limited product is still an important change. Meta has spent years showing realistic avatars as research, usually in demonstrations associated with VR. Hologram moves the idea into a consumer communication service with a defined platform, market, and launch window.

The move also separates the product from its name. Nothing is projected into physical space for the person receiving a Ray-Ban Display call. On that device, a Hologram is an AI-generated video portrait designed to substitute for a missing selfie camera.

Why Meta Put Its Realistic Avatar in Smart Glasses First

Ray-Ban Display gives Meta a constrained problem that generative video can solve before full spatial telepresence is ready.

Meta’s realistic-avatar research has traditionally aimed higher. Its Codec Avatars work explored photorealistic people who could occupy shared virtual environments. The goal was not merely a better profile picture, but convincing eye contact, expression, and physical presence across distance.

A 2020 paper on Modular Codec Avatars described faces driven by cameras mounted on a VR headset. That approach pursued a three-dimensional representation connected to tracked behavior.

Hologram’s first consumer version changes the mechanism. It does not require inward-facing cameras to continuously observe the wearer’s face. It generates a two-dimensional output from audio and a previously captured identity.

That compromise fits Meta Ray-Ban Display. Glasses intended for regular public use must remain relatively compact and socially familiar. Filling their frame with extra facial-tracking cameras would create different hardware, power, and design constraints.

Voice is already available during a call. Meta can therefore use the microphones as a behavioral input and move the expensive generation work to its servers. The glasses collect less facial information during the conversation because they cannot directly see the wearer’s face.

This approach also gives Meta an immediate answer to a product weakness. Existing glasses calls can share the wearer’s viewpoint, which works well for showing a scene or asking for remote assistance. They are less suitable when the other participant wants ordinary face-to-face contact.

Meta illustrates the gap with a cooking call. Someone wearing the glasses can continue preparing food while a family member sees the person’s digital face. The caller avoids holding or positioning a phone, and the recipient gets something closer to a conventional video conversation.

The scenario captures both the appeal and the uncertainty. Hands-free communication is useful while cooking, repairing equipment, walking through a workspace, or caring for a child. However, a fixed synthetic background can hide the context that makes those moments meaningful.

The generated face also does not necessarily show the caller’s literal expression. Tone and language provide strong signals, but they cannot capture every glance, distraction, suppressed reaction, or physical event. A convincing output is therefore an interpretation of the user, not a transparent window onto the user.

Meta’s research helps explain why it believes audio can carry more behavior than a conventional avatar system suggests. In 2025, the company described conversational motion models trained to connect audiovisual signals with facial expressions, gestures, and listening behavior.

That research involved more than 4,000 hours of two-person interactions and over 4,000 participants. It supports a broader effort to model how speech, movement, and social responses fit together.

Hologram applies the same general idea to a narrower consumer experience. Rather than reconstructing only what sensors directly measure, the system predicts an expressive performance that looks plausible.

This is why smart glasses come first. A phone already has a front camera, so replacing its live image would offer fewer practical benefits. A VR headset can support richer tracking, but it also raises the technical bar for true spatial presence.

Ray-Ban Display sits between those devices. It creates a real video-calling limitation while allowing Meta to ship a useful-looking version before its larger avatar vision is ready.

Meta Holograms Take a Different Route From Apple Personas

The central contest is not simply Meta versus Apple. It is prediction from sparse signals versus reconstruction from dedicated headset sensors.

Apple’s Persona system provides the clearest comparison because it also lets a person appear on calls while a device covers the face. Apple introduced Personas with Vision Pro, then improved their realism and expression through later visionOS releases.

A Vision Pro user creates a Persona by scanning their appearance with the headset. During a call, the device tracks facial and hand behavior through its sensor system. Apple describes a Persona as a spatial representation that conveys the wearer’s expressions and movements in real time.

Personas also work beyond FaceTime. Apple said at launch that supported applications included Zoom, Cisco Webex, and Microsoft Teams. This positioned the feature as a communication layer for a spatial-computing headset rather than a single service effect.

Meta Holograms address the same basic obstruction problem but start from different hardware. Meta Ray-Ban Display does not cover the face like Vision Pro. It simply lacks a camera pointed back at the wearer.

Apple can use a sensor-rich headset to observe behavior hidden behind the device. Meta instead asks a server-side model to infer behavior that its glasses cannot observe directly.

The difference becomes clearer through setup and output:

Capture and tracking

  • Meta: A phone captures identity, expressions, clothing, and background during setup. Microphone audio drives the fall release during calls.

  • Apple: Vision Pro creates the Persona and uses headset sensors to animate expressions and movement during use.

Initial destination

  • Meta: A shoulders-up video representation appears in WhatsApp calls received on ordinary screens.

  • Apple: A spatial representation appears in FaceTime and supported conferencing applications.

Hardware context

  • Meta: The caller wears lighter, socially familiar display glasses designed for use across everyday activities.

  • Apple: The caller wears an immersive spatial computer with extensive tracking hardware.

Product tradeoff

  • Meta: Lower-friction hardware requires more behavioral inference and cloud generation.

  • Apple: Richer sensing supports more directly tracked behavior, but requires a substantially different device category.

Apple has continued improving the visual result. Its visionOS 26 preview highlighted sharper profiles, more natural hair, eyelashes, and complexions. It also added improved setup controls and shared spatial experiences.

Those refinements matter because digital likenesses face an unusually demanding quality threshold. Minor problems with eyes, timing, gaze, or skin movement can attract more attention than large errors in a cartoon avatar.

Meta’s solution attempts to cross that threshold by generating a familiar video format. In an UploadVR hands-on, David Heaney initially mistook a Meta employee’s Hologram for a normal selfie-camera feed. He described the output as plausible, while also questioning its usefulness on an ordinary phone screen.

That reaction identifies Meta’s opportunity. Hologram does not need to reproduce every movement perfectly to work as a casual call presence. It needs to look believable enough at phone-screen size while staying synchronized with speech.

It also identifies the risk. A result that looks like a normal camera feed can obscure the fact that it is synthetic. Apple’s Persona has often looked visibly distinct from conventional video, making the mediation easier to recognize.

Meta must therefore solve more than visual fidelity. It needs a clear interaction language for generated identity. Recipients should know when a real camera view ends and an inferred performance begins.

A Plausible Face Is Not the Same as a Faithful Face

Hologram’s main technical achievement creates its hardest trust problem: the better it looks, the easier it becomes to mistake inference for observation.

A normal video call has many flaws. Cameras crop the scene, networks drop frames, lighting distorts skin, and participants choose what enters the frame. Yet the basic expectation remains that the camera records light from the person and surroundings at that moment.

Meta Holograms change that relationship. The output depicts the caller, but its immediate behavioral input is audio. The model decides how speech should look on the captured identity.

That decision can work well during ordinary conversation. Speech carries timing, emotional tone, emphasis, pauses, and other expressive cues. A trained model can turn those signals into convincing lip movements and facial responses.

However, voice cannot fully specify the state of a face. A person can smile while delivering bad news, glance away during an uncomfortable question, or silently react to something outside the call. The fall version lacks a direct view of those moments.

The Hologram can therefore appear more attentive, expressive, or composed than the caller. It may show a statistically fitting reaction instead of the wearer’s actual one.

This distinction matters differently across situations. During a casual family call, participants might accept the avatar as a convenient stand-in. During a sensitive workplace meeting, medical conversation, interview, or negotiation, inferred expression can alter how a message is received.

Meta has not publicly detailed every disclosure, consent, retention, or moderation control associated with Hologram. The available announcement confirms server-side generation and a personalized setup, but it does not answer every question raised by that architecture.

Users will need plain explanations of what leaves the device, how long capture assets persist, and whether they can delete or regenerate their model. Call recipients also need an obvious signal that the image is synthetic.

Identity protection presents another unresolved area. A personalized model capable of producing realistic video from voice becomes sensitive material. Meta needs safeguards against unauthorized activation, copied capture data, account compromise, and deceptive reuse outside the intended call flow.

The system also depends on remote computation. That creates questions about latency, service interruptions, network requirements, and regional availability. Meta’s Early Access restriction to the United States suggests a deliberately controlled first release.

The captured clothing and static background create less serious but more immediate social problems. A caller might appear in yesterday’s shirt, in a room they no longer occupy, or against a background that suggests a false location.

Those mismatches could remind recipients that they are seeing an avatar. They could also create confusion if Hologram resembles a live feed closely enough that people assume the scene is current.

Visual inclusion deserves attention as well. Photorealistic systems must represent different faces, skin tones, hair types, accessories, speech patterns, and expressions consistently. A model that performs unevenly can make some users look less accurate or less expressive than others.

Early Access should generate evidence on these issues, but Meta has not published broad independent performance measurements for the consumer feature. A successful stage demonstration establishes feasibility, not reliability across millions of people and uncontrolled conditions.

The right standard is not whether Hologram can fool someone briefly. It is whether the system remains useful after participants understand exactly how it works.

Trust will depend on consistency, disclosure, user control, and appropriate expectations. Realism alone cannot supply those qualities.

Full-Body Holograms Reveal How Much Work Remains

Meta’s next Hologram looks more spatial, but its launch architecture still falls short of the company’s original co-presence ambition.

Meta plans a full-body version for its VR Glasses in spring 2027. When two users call through WhatsApp, they will first see Holograms inside a three-dimensional window. A control can then place the other person at approximately life size in the viewer’s space.

This version uses more tracking information than the Ray-Ban Display release. It also creates a stereoscopic stream, meaning each eye receives a slightly different image to produce a sense of depth.

Yet the launch product is not a volumetric person. It does not construct an avatar that the viewer can inspect naturally from every side. Instead, it presents a masked stereoscopic video plane that keeps facing the observer.

The technique can look compelling while both participants remain in favorable positions. UploadVR’s demonstration found the result effective while standing still or sitting. The sense of scale and depth made a remote person feel closer to occupying the room.

Movement exposed the compromise. When the observer leaned or walked sideways, the Hologram rotated to maintain its front-facing view. The behavior resembled a billboard sprite in a video game rather than a body with stable three-dimensional geometry.

Fine details also appeared blocky during the demonstration, particularly around the eyes. Meta engineers reportedly pointed to congested event Wi-Fi as one possible factor, but the public test did not isolate the cause.

These limitations separate Hologram from Meta’s earlier Codec Avatar promise. Research prototypes aimed to reconstruct and animate a person as a true three-dimensional presence. The first shipping versions instead generate video streams optimized for particular viewing conditions.

That is the article’s central reversal. Meta is finally preparing to ship realistic avatars after years of research, but it is doing so by narrowing the problem. The company is not placing its most ambitious laboratory system directly into consumer hardware.

The shortcut is commercially understandable. A high-quality stream can deliver much of the visual impact without solving every problem of geometry, relighting, gaze, capture, and real-time rendering.

It may also help Meta collect evidence about what people value. Users might care more about recognizable faces and natural conversation than walking around a perfect volumetric reconstruction. If so, video generation could become the practical foundation rather than a temporary substitute.

Meta says a later Hologram version will become truly volumetric, support group conversations, and become available to outside developers. No specific timetable accompanies that commitment.

The distinction matters for developers. A full spatial avatar could interact with application geometry, occupy stable positions, and support experiences beyond calls. A front-facing stream offers far less flexibility.

Meta’s current Avatar SDK, built around stylized characters, supports experiences where consistency and animation matter more than exact likeness. Replacing it with realistic representations would require predictable performance, permissions, integration tools, and behavior across many applications.

Hologram has not reached that stage. The 2027 version should be judged as a telepresence stream designed for Meta hardware and WhatsApp, not as the completion of the company’s broader avatar platform.

Three Signals Will Decide Whether Hologram Becomes More Than a Demo

The next test is not another controlled presentation. It is whether Meta can turn visual plausibility into sustained, trusted communication.

The first signal is the US Early Access rollout on Meta Ray-Ban Display. Meta needs to ship within its stated fall window and show that setup, call initiation, generation, and delivery work outside Connect.

Reliability will matter as much as image quality. Users must be able to create a stable likeness without repeated scanning, and the generated feed must remain synchronized under normal network conditions.

Recipient behavior offers the more revealing adoption measure. People must prefer receiving a Hologram over continuing with an audio call or seeing the wearer’s outward camera. If the feature feels unnecessary, awkward, or misleading, technical realism will not rescue it.

Meta should also make synthetic status clear inside WhatsApp. A recognizable indicator, combined with accessible controls and explanations, would strengthen the case that Hologram is a communication mode rather than an imitation camera feed.

The second signal is movement input. Meta has said it plans to incorporate the glasses’ IMU so real head rotation can influence the avatar. That change would reduce the gap between inferred and observed behavior without adding inward-facing cameras.

The update would also reveal Meta’s long-term balance between sensors and generation. Each measured input improves fidelity but adds data, engineering, and potentially power requirements. Each predicted movement preserves simplicity but increases the chance of inaccurate expression.

The third signal is the spring 2027 full-body launch. Meta must show that the stereoscopic approach works consistently across home networks and varied physical spaces. It should also clarify when, or under what conditions, Hologram will become genuinely volumetric.

Apple’s response will provide useful context. Continued Persona improvements could strengthen the argument for sensor-rich reconstruction. Meta’s progress could support the competing idea that generative models can produce convincing presence from fewer inputs and more familiar hardware.

For developers and enterprise buyers, the immediate lesson is caution. Hologram is not yet a general avatar service, developer platform, or substitute for verified visual presence. It is a narrowly scoped WhatsApp feature entering a limited rollout.

For consumers, the choice will be more personal. A synthetic face can make hands-free glasses calls warmer and more familiar. It can also replace truthful visual detail with a model’s best guess.

Meta Holograms will succeed only if that exchange feels worthwhile after the novelty fades. Watch whether users keep the feature enabled, whether recipients trust its presentation, and whether Meta connects more real movement to the generated face.

The defining question is simple: when someone calls from smart glasses, do you want to see what a model thinks they look like, or would you rather hear their voice and know what remains off camera?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page