Perceptron: Perceptron Mk1.5
About this model
Perceptron Mk1.5 is a closed-source multimodal model from Perceptron AI, released in September 2026 as the flagship offering in the Perceptron family for embodied and physical-world agents. It accepts text, images, video, and audio through the Perceptron API (model ID perceptron-mk1.5) and responds with natural language plus machine-readable spatial and temporal structure: points, bounding boxes, polygons, video clips, and timestamped object tracks. The design targets deployments on drones, quadruped robots, smart glasses, and mobile devices where downstream systems need geometry and identity over time rather than captions alone.
Beyond perception, Mk1.5 supports configurable reasoning effort, OpenAI-compatible function calling, JSON Schema constrained outputs, native video object tracking, audio-visual understanding (including optional video soundtracks), web search and other user-defined tools, and parallel sub-agent calls. Perceptron reports strong results on video tracking, egocentric hand localization, and multimodal search with tools, alongside competitive text-centric scores on MMLU-Pro, GPQA-Diamond, and LiveCodeBench in its public performance materials. Inference is served via the Perceptron Platform and SDK; weights are not released for local hosting.
Official API documentation specifies a 36,864-token context window (the launch post describes 32K tokens of multimodal context), up to 8,192 output tokens, and pricing of $0.15 per million input tokens and $1.50 per million output tokens. Parameter count and full training architecture are not disclosed publicly.
Benchmark Scores
Technical Specs
- Architecture: Transformer
- Context Window: 36,864 tokens
- Input Modalities: text, image, audio
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.15 / 1M tokens | 1.50 / 1M tokens | USD |