Neural response analytics for video
Know how a video lands before you pay to run it.
Bayan predicts how the brain responds to your video, second by second. You see where attention holds, where it drifts, and what people are likely to feel, so you can fix the weak moments before the budget goes out.
In this demo your video is processed in your browser and never uploaded.
left hemisphere
What the report shows
Each measure is computed for every quarter second of the video, then summarized. Click any moment in the report to jump the video there.
- Attention 0–100, per 0.25 s
- How strongly the dorsal attention network is engaged relative to the default-mode network.
- Hook first 3 s
- Whether the opening earns the next ten seconds. On a vertical feed this decides most of your views.
- Brain response z-score
- Predicted activity on the cortical surface and in MNI slices, from visual and auditory cortex to language and face areas.
- Emotional arc 7 emotions
- Interest, joy, surprise, calm, tension, sadness and boredom over time, plus valence and arousal.
- Predicted recall 0–100
- How likely the key moments are to be remembered, based on predicted memory-encoding activity.
- Moments to fix timestamped
- Where attention drops, why it probably drops, and what to change in the edit.
How it works
Upload a cut
Drop in an MP4, MOV or WebM. Add the campaign goal and where it will run so the recommendations fit the placement.
The model watches it
It reads the picture, the sound and the faces on screen, then predicts the response of each brain region at every moment.
Read the report and edit
Scores, a brain view, an emotion timeline and a short list of timestamped fixes. Re-run the new cut and compare.
Built for teams that pay for every view
Communications and development teams at nonprofits spend real money putting video in front of donors and supporters. Most of them can't afford to A/B test every cut on paid social. Bayan gives them a read on a cut before it goes live.
Find the moment with the most attention and feeling, and put the ask there.
Check that the key fact lands while people are still watching.
See whether hard footage drives engagement or pushes people away.
About the model
The production model is built on TRIBE v2, an open brain encoding model trained on fMRI recordings of people watching films. Given a video, it predicts activity at each of the 20,484 points of the fsaverage5 cortical surface over time.
This demo runs a lightweight approximation in your browser so anyone can try it without an account or a GPU. It measures the video's visual, audio and facial signals and maps them onto the same cortical surface and the same brain networks. Treat its numbers as directional, not as a substitute for the full model.
Questions
Does my video get uploaded anywhere?
No. In this demo, the frames, audio and face detection are all processed in your browser tab. Closing the tab clears the video.
Which files work?
MP4 (H.264), WebM and MOV. iPhone videos are usually HEVC, which Safari and most Macs decode fine. If your browser can't play the file, export it as MP4 (H.264) and try again. Videos up to 5 minutes work best.
How accurate is it?
The full model predicts group-average brain responses measured with fMRI. It tells you how an average viewer's brain is likely to respond, not how a specific person will act. This demo is a simplified version of that, so use it to compare cuts and spot weak moments rather than to forecast exact results.
Does this replace A/B testing?
It comes before it. Use it to pick the strongest two or three cuts, then spend your test budget on those instead of on every version.
What do the colors on the brain mean?
They follow fMRI convention. Red to yellow means predicted activity above baseline, and blue means below. The numbers are z-scores, so +3 is a strong response.
New analysis
Analyze a video
Add a cut and tell us where it will run. A 30-second video takes about a minute.
- Load video and models
- Decode audio
- Sample frames, detect faces
- Predict brain response
- Score and write report
Report
Report
Most active now
z-score
Slices
MNI152, 2 mm. Click a slice to move the crosshair.
Network response over time
Warm above baseline, blue below
Regions
Click a row to jump to its peak
| Region | Responds to | Mean z | Peak z | Peak at |
|---|
Right now
Valence and arousal
Path through the video
Emotional arc
Predicted intensity, 0 to 1
Share of running time
By strongest emotion
Faces on screen
Timeline
Attention with the strongest emotion underneath. Click or drag to scrub.
Key moments
What to change
Fix first
Working
Method and limits of this demo
What is measured. Frames are sampled every 0.25 s (or 0.5 s in fast mode). For each frame the demo measures contrast, edge density, saturation, warmth, motion, shot changes and on-screen text. The audio track is measured for loudness, speech-band energy, spectral change and pitch. A small face model (TinyFaceDetector with a facial-expression classifier) runs every half second.
How it maps to the brain. These signals drive 16 response channels, such as early visual, motion, face, language, attention, salience and default mode. Each channel is placed on the cortex using 300 network ROIs from Seitzman et al. (2020) plus named regions like the fusiform face area and Broca's area. Activity spreads from each ROI with a Gaussian kernel on the fsaverage5 surface and the MNI152 volume.
What it is not. This is not the TRIBE v2 network. The full model learns these mappings from fMRI data and captures things this demo can't, like the meaning of what is said. The demo is useful for spotting flat stretches and comparing cuts. It should not be quoted as measured brain data.
Timing. Real BOLD signals lag the stimulus by about 5 seconds. The report removes that lag so each value lines up with the frame that caused it.