Video and embodied reasoning
Mk1 is described as purpose-built for video understanding and embodied reasoning, with an emphasis on temporal reasoning across continuous streams instead of isolated snapshots.
Perceptron Mk1 is a closed-source vision model for video understanding and embodied reasoning, with API access and structured outputs for robotics and physical-world workflows.
Perceptron Mk1 is a closed-source model from Perceptron built for video understanding and embodied reasoning. The company describes it as a layer of intelligence for the physical world, aimed at workloads where perception, timing, and spatial grounding matter more than text-only generation.
The model is positioned for physical AI and robotics workflows, with support for image, video, and embodied reasoning, plus structured outputs such as points, boxes, polygons, tracks, clips, HTML, JSON, and Markdown. The source pages also show developer examples for detection, pointing, counting, OCR, captioning, and promptable visual analysis through APIs.
Mk1 is described as purpose-built for video understanding and embodied reasoning, with an emphasis on temporal reasoning across continuous streams instead of isolated snapshots.
The model can reason over time, produce structured breakdowns of events, and optionally turn reasoning off when it is not needed.
It analyzes video at a dynamic frame rate of up to 2 FPS in a 32K-token context window and can return structured timecodes for specific moments.
The site says one reference image or video can be used to find matching instances across new media, and two pieces of media can be compared without fine-tuning or a labeled dataset.
Mk1 supports pointing, counting, OCR, document extraction, and other image-reasoning tasks, including messy text, analog gauges, and tables with preserved structure.
The model is trained to emit spatial primitives such as point, box, polygon, track, and clip, which can be consumed directly by downstream systems.
Use Mk1 to interpret teleoperation footage, label subtask boundaries, extract success or failure signals, and turn raw episodes into supervised data for downstream policy training.
Apply the model during inference to return grasp affordances, constraint checks, relational targets, and cross-camera tracking for manipulation or navigation systems.
Run the model on factory, warehouse, or construction imagery and video to detect defects, flag safety issues, and read instruments during inspection rounds.
Use temporal grounding and structured outputs to clip sports moments, search film and TV libraries, or moderate AI-generated content at scale.
Analyze satellite, drone, and fixed-camera footage for infrastructure monitoring, construction progress, vegetation encroachment, or post-disaster assessment.
Perceptron Mk1 is built for video understanding and embodied reasoning, with additional support for image reasoning and structured document extraction. The site positions it for physical-world applications rather than general chat.
The developer page shows Python-style examples for tasks such as focus/zoom and crop, conversational pointing, in-context learning, object detection, counting, OCR, and captioning. The demo also shows a mode for segmenting one or multiple classes in an image.
The site says Mk1 analyzes video at up to 2 FPS within a 32K-token context window and can return structured timecodes, clips, and other spatial outputs such as points, boxes, polygons, tracks, and clips.
The homepage says Mk1 is a closed-source model family release. The site also says developers can use the model through APIs or reach out for a commercial license to the weights.
The pricing page does not show published plan details on the collected text, so exact pricing, tiers, and limits are not available from the source pages used here.
AakarDev AI helps teams manage AI provider access, project setup, logs, and analytics in one dashboard. BYOK support included.
Arduino VENTUNO Q is an edge AI computer for AI and robotics, combining AI inference and deterministic control on one board with Arduino App Lab support.
Benchspan is an AI agent security platform that discovers agents, blocks prompt injection and data exfiltration in real time, and supports pre-launch red teaming.
Edgee is an AI gateway for coding agents and LLM apps. It compresses tokens, routes requests across models, and adds observability and team controls.
CreateOS Sandbox is an isolated compute environment for running code and agent workloads in Firecracker micro-VMs with private networking and SDK, CLI, or MCP control.
Codex Plugins bundle reusable skills, app integrations, and MCP servers into workflows you can install in the Codex app or use from Codex CLI.