JARVIS
A wearable AI, built as one system.


Wearable systems, EEG models and spatial intelligence. From the first circuit to a working agent.
Design visualization. Actual prototype photographs follow below.
Hardware you can see. Research you can inspect.
A wearable AI, built as one system.
+6.75 ppBETA, 1.2 s vs. protocol-adapted SSVEPformer
Public-data study. Selection-aware comparison.
1,281recorded keyframesWatch the field demo
I build systems that connect perception, memory and action. JARVIS is the working foundation; EEG and spatial computing extend how people can interact with it.
A glasses camera, local phone inference and a memory-based agent form one working system.
Time-aligned vision, speech, location and message summaries feed personal memory. Habit-triggered semantic review decides when an action or reminder is useful.
Running measurements supplied by the builder, October 2026. Not independently re-benchmarked for this portfolio. Single-frame inference only, excluding transport, retrieval and decisions.
Measurement scopeExplore the path from a camera frame to a contextual action.
Architecture walkthrough, not a live device connection.
The phone runs local visual inference with QNN/HTP. Frame IDs align descriptions with visual-token identity results. Temporal checks retain evidence across observations.
Summary and retrieval agents reduce context load. Local bilingual embeddings, lexical retrieval, reranking and time/source metadata support contextual recall and object search.
Generated habit state machines provide a fast trigger. The main agent reviews the complete observation before intervening. Tool completion relies on execution evidence and device acknowledgements.
Camera/audio transport, glasses UI, Android local VLM integration, temporal visual reasoning, memory retrieval, real-time speech context and tool execution are implemented. Foundation models and speech services are dependencies. The RV1106B prototype uses a development kit; custom hardware development is ongoing.
Visual-token matching reuses features already produced by the local VLM. It supports learned places, negative examples, region records and same-frame identity supplementation. Appearance matching and object ownership still need scenario-specific validation.
Foldable convolution, harmonic-biased attention and a compact Transformer for cross-subject EEG decoding.
Compare observation windows and public datasets. Each neural result averages participant-disjoint evaluation over three training seeds.
Explore how the local mixer changes at deployment. The reduced-token Transformer and harmonic attention remain in the model.
Schematic, not an attention heatmap. Equivalence checks cover the retained checkpoints and tested windows.
Foldable local temporal mixing precedes time-domain and complex-spectral candidate evidence. Reduced spectral tokens feed harmonic-biased late attention and Transformer candidate modeling. An optional adapter updates just 113 parameters for participant calibration.
In the registered BETA ablation, including local temporal mixing improves accuracy by 3.99-7.23 percentage points across the tested windows. Including late attention lowers it by 0.66 points at 0.4 s and improves it by 1.20-2.16 points at 0.6-1.5 s.
All four datasets are public. Benchmark and BETA are selection-aware because later candidate screening consulted them. Wearable dry/wet analysis was frozen after model retention. The model is not the best at 0.4 s; ordinary SSVEPformer also runs faster on the measured laptop.
CPU times use batch one, float32 and one inference thread, with 100 warm-up and 1,000 measured runs. They include registered model preprocessing, exclude signal acquisition, and do not establish online or glasses latency. No final glasses-mounted electrode montage has been validated.
| Window | HarmonicFoldNet | SSVEPformer |
|---|---|---|
| 0.4 s | 27.37% | 30.26% |
| 0.6 s | 43.47% | 43.46% |
| 0.8 s | 55.60% | 52.79% |
| 1.0 s | 63.56% | 58.52% |
| 1.2 s | 68.70% | 61.95% |
| 1.5 s | 73.12% | 65.09% |
Visual recognition answers where. A persistent spatial model can also describe how places, objects and people relate.
The scenic navigation prototype combines recorded maps, visual relocalization, route instructions and glasses display. The recorded route below contains 1,281 keyframes.
Recorded route playback, not live localization. Monocular map coordinates are not metre-scale distances.

The next research step explores scene reconstruction with Gaussian splatting, stable spatial anchors, occlusion and spatially grounded AI reasoning. The aim is to place assistance in the surrounding environment and let memory refer back to it.
Gaussian reconstruction is an active research branch. A complete validated spatial-computing integration is still ahead.
EEG, temporal AI and spatial computing are parallel directions with distinct validation steps.
Compare posterior scalp baselines with peri-auricular and temple recordings. Test repeat wearing, drift, motion artifacts, personal calibration and no-command rejection before selecting a practical montage.
Compare distinct locations TP9 and T9, TP10 and T10, FT9 and F9, FT10 and F10. Nz is optional; O1, O2 and Oz remain posterior controls. These are candidate comparisons, not interchangeable site names or a validated glasses design.
Earlier public-data scalp/ear experiments used 12 participants and 1,440 unique trials. The 24-channel scalp study and ear study share the same participants. They form a separate montage feasibility study, not an additional 24-person cohort.
Behind-the-ears SSVEP studyMobile scalp and ear EEG datasetTrain models that connect successive observations, object changes and user actions. Evaluate earlier intent recognition, coherent memory and timely intervention under compute and false-trigger budgets.
Combine scene geometry, semantic objects and persistent memory. Test stable world anchors, incremental map reuse and object retrieval that can guide a user back to the last supported location.
Evaluate deliberate silent articulation with surface EMG or tissue-conducted sound as a separate future input study. These signals require their own hardware and controlled data; existing EEG results do not establish this capability.
Different contributions, with the implementation and evidence made explicit.
Foldable local-to-global EEG decoder, multi-dataset evaluation, three seeds, ablations, participant adaptation and deployment equivalence.
Visual-to-language connector training and LoRA experiments. Saved checkpoints and a completed ten-epoch second-stage log document the training work.
Teacher-guided visual distillation, projector alignment and generation smoke tests. The existing evidence supports a training pipeline, not a finished large-scale pretrained model.
QNN/HTP integration on Snapdragon 8 Gen 3. Place and identity matching reuse local visual features, with frame alignment and unknown-rejection research.
Code and model releases offer a way to examine the work beyond this page.
Source-available, noncommercial research release. Code, weights, protocols and derived source data.
A scenario sandbox with selectively activated agents and bounded model-call budgets. Exploratory results, not forecasts.
Build and C++/CUDA compatibility work for an existing SLAM system on sm_120 GPUs.
A connectome-constrained LIF simulation constructed from official male fly CNS data, with 166,700 annotated neurons and 25,582,938 directed connections. Stimulus, no-stimulus and sign-policy sensitivity conditions were tested.
This is a computational research branch. Synapse counts are not measured conductances, and the simulation is not yet a validated EEG decoder.
Official MaleCNS sourceThe first experiments were small. Display optics, motion sensing and simple mechanisms led toward the present wearable system.
ESP32, OLED and a prism formed an early wearable display. A later version used a camera viewfinder as its display source.
ESP32 and MPU6050 connected motion sensing with light control.
An enclosure, clamp and torque motor tested controlled belt contraction.
Hardware development, local visual inference, memory agents and neural-interface research now share one engineering direction.