XzStark
Polished graphite wearable glasses design visualization in the original front three-quarter viewPolished graphite wearable glasses design visualization in the original front three-quarter view
Independent builder

I build AI
you can wear.

Wearable systems, EEG models and spatial intelligence. From the first circuit to a working agent.

Design visualization. Actual prototype photographs follow below.

What I've built.

Hardware you can see. Research you can inspect.

One direction.
Three ways in.

I build systems that connect perception, memory and action. JARVIS is the working foundation; EEG and spatial computing extend how people can interact with it.

JARVIS. An AI that shares your world.

A glasses camera, local phone inference and a memory-based agent form one working system.

Actual near-eye display, photographed through the prototype optics.

Perception becomes context.

Time-aligned vision, speech, location and message summaries feed personal memory. Habit-triggered semantic review decides when an action or reminder is useful.

Snapdragon 8 Gen 3

~600ms
Fastest single-frame result
~1s
Average single-frame inference

NVIDIA RTX 5060

~300ms
Fastest single-frame result
~600ms
Average single-frame inference

Running measurements supplied by the builder, October 2026. Not independently re-benchmarked for this portfolio. Single-frame inference only, excluding transport, retrieval and decisions.

Measurement scope

One shared context.
Specialized agents.

Explore the path from a camera frame to a contextual action.

GlassesCamera & audio
Local perceptionQNN / HTP
Shared memoryRetrieve & summarize
Reviewed actionEvidence & ACK

Architecture walkthrough, not a live device connection.

Glasses cameraPhone VLMTemporal evidence

The phone runs local visual inference with QNN/HTP. Frame IDs align descriptions with visual-token identity results. Temporal checks retain evidence across observations.

Implementation and contribution

Camera/audio transport, glasses UI, Android local VLM integration, temporal visual reasoning, memory retrieval, real-time speech context and tool execution are implemented. Foundation models and speech services are dependencies. The RV1106B prototype uses a development kit; custom hardware development is ongoing.

Visual-token matching reuses features already produced by the local VLM. It supports learned places, negative examples, region records and same-frame identity supplementation. Appearance matching and object ownership still need scenario-specific validation.

HarmonicFoldNet-SSVEP

Foldable convolution, harmonic-biased attention and a compact Transformer for cross-subject EEG decoding.

247unique public-data participants
435kfolded executable parameters
3.04msmedian single-thread CPU inference at 1.2 s
49,000folding checks, zero label changes

Evidence you
can explore.

Compare observation windows and public datasets. Each neural result averages participant-disjoint evaluation over three training seeds.

1.2 s68.70%HarmonicFoldNet balanced accuracy+6.75 ppversus protocol-adapted SSVEPformer
BETA: participant-mean balanced accuracy by observation windowHarmonicFoldNet compared with a protocol-adapted SSVEPformer. Observation windows from 0.4 to 1.5 seconds. Source values are available in the data table.0204060801000.40.60.81.01.21.5Observation window (seconds)
HarmonicFoldNetSSVEPformer
70 participants

Local branches merge.
Attention stays.

Explore how the local mixer changes at deployment. The reduced-token Transformer and harmonic attention remain in the model.

EEG
ConvIdentityNormΣ
Merged local operator
Reduced tokensHarmonic attention
Shared
scorer
435,043deployment parameters
49,000 / 0folding checks / label changes

Schematic, not an attention heatmap. Equivalence checks cover the retained checkpoints and tested windows.

Benchmark35 participants
BETA70 participants
Wearable102 participants / dry & wet
Kim2025BetaRange40 participants
Architecture, adaptation and module contributions

Foldable local temporal mixing precedes time-domain and complex-spectral candidate evidence. Reduced spectral tokens feed harmonic-biased late attention and Transformer candidate modeling. An optional adapter updates just 113 parameters for participant calibration.

Original architecture figure from the reproducibility repository.

In the registered BETA ablation, including local temporal mixing improves accuracy by 3.99-7.23 percentage points across the tested windows. Including late attention lowers it by 0.66 points at 0.4 s and improves it by 1.20-2.16 points at 0.6-1.5 s.

Original attention analysis figure from the research repository.
Evaluation boundaries and reproducible values

All four datasets are public. Benchmark and BETA are selection-aware because later candidate screening consulted them. Wearable dry/wet analysis was frozen after model retention. The model is not the best at 0.4 s; ordinary SSVEPformer also runs faster on the measured laptop.

CPU times use batch one, float32 and one inference thread, with 100 warm-up and 1,000 measured runs. They include registered model preprocessing, exclude signal acquisition, and do not establish online or glasses latency. No final glasses-mounted electrode montage has been validated.

BETA participant-mean balanced accuracy
WindowHarmonicFoldNetSSVEPformer
0.4 s27.37%30.26%
0.6 s43.47%43.46%
0.8 s55.60%52.79%
1.0 s63.56%58.52%
1.2 s68.70%61.95%
1.5 s73.12%65.09%
Download evidence snapshot

AI with a sense of space.

Visual recognition answers where. A persistent spatial model can also describe how places, objects and people relate.

From recognition
to spatial memory.

The scenic navigation prototype combines recorded maps, visual relocalization, route instructions and glasses display. The recorded route below contains 1,281 keyframes.

0 / 1,281

Recorded route playback, not live localization. Monocular map coordinates are not metre-scale distances.

Recorded route connecting Mingzhi Square and Ningxia Luoying Pavilion
Mingzhi SquareNingxia Luoying Pavilion

Gaussian splatting
and spatial AI.

The next research step explores scene reconstruction with Gaussian splatting, stable spatial anchors, occlusion and spatially grounded AI reasoning. The aim is to place assistance in the surrounding environment and let memory refer back to it.

Gaussian reconstruction is an active research branch. A complete validated spatial-computing integration is still ahead.

What I'm working toward.

EEG, temporal AI and spatial computing are parallel directions with distinct validation steps.

Glasses-compatible EEG

Compare posterior scalp baselines with peri-auricular and temple recordings. Test repeat wearing, drift, motion artifacts, personal calibration and no-command rejection before selecting a practical montage.

Candidate locations and first experiments

Compare distinct locations TP9 and T9, TP10 and T10, FT9 and F9, FT10 and F10. Nz is optional; O1, O2 and Oz remain posterior controls. These are candidate comparisons, not interchangeable site names or a validated glasses design.

Earlier public-data scalp/ear experiments used 12 participants and 1,440 unique trials. The 24-channel scalp study and ear study share the same participants. They form a separate montage feasibility study, not an additional 24-person cohort.

Behind-the-ears SSVEP studyMobile scalp and ear EEG dataset

Temporal perception models

Train models that connect successive observations, object changes and user actions. Evaluate earlier intent recognition, coherent memory and timely intervention under compute and false-trigger budgets.

Spatially grounded assistance

Combine scene geometry, semantic objects and persistent memory. Test stable world anchors, incremental map reuse and object retrieval that can guide a user back to the last supported location.

Optional silent interaction

Evaluate deliberate silent articulation with surface EMG or tissue-conducted sound as a separate future input study. These signals require their own hardware and controlled data; existing EEG results do not establish this capability.

Models built, adapted and deployed.

Different contributions, with the implementation and evidence made explicit.

HarmonicFoldNet-SSVEP

Architecture & training

Foldable local-to-global EEG decoder, multi-dataset evaluation, three seeds, ablations, participant adaptation and deployment equivalence.

TinyVLM / LoRA

Alignment & fine-tuning

Visual-to-language connector training and LoRA experiments. Saved checkpoints and a completed ten-epoch second-stage log document the training work.

Visual encoder student

Distillation pipeline

Teacher-guided visual distillation, projector alignment and generation smoke tests. The existing evidence supports a training pipeline, not a finished large-scale pretrained model.

On-device VLM & token reuse

Runtime integration

QNN/HTP integration on Snapdragon 8 Gen 3. Place and identity matching reuse local visual features, with frame alignment and unknown-rejection research.

Public work, open to inspection.

Code and model releases offer a way to examine the work beyond this page.

Computational neuroscience: MaleCNS

A connectome-constrained LIF simulation constructed from official male fly CNS data, with 166,700 annotated neurons and 25,582,938 directed connections. Stimulus, no-stimulus and sign-policy sensitivity conditions were tested.

This is a computational research branch. Synapse counts are not measured conductances, and the simulation is not yet a validated EEG decoder.

Official MaleCNS source

A longer
practice of building.

The first experiments were small. Display optics, motion sensing and simple mechanisms led toward the present wearable system.

Original 2022 display experiment.
2022

Reflection display

ESP32, OLED and a prism formed an early wearable display. A later version used a camera viewfinder as its display source.

Earlier work

Gesture lighting

ESP32 and MPU6050 connected motion sensing with light control.

Motorized belt mechanism

An enclosure, clamp and torque motor tested controlled belt contraction.

Current work

Devices, models and systems

Hardware development, local visual inference, memory agents and neural-interface research now share one engineering direction.