AI News Feed
Market watch
Category

Computer Vision

Latest news, developments and reporting on Computer Vision.

Computer Vision · Large Language Models

Chinese AI Startup Shows Similar World Model Six Months Ahead of Atlas, Tops Benchmark

World Labs released Atlas on Sept. 2, but Chinese startup Yingsu had already open-sourced a comparable world model in March and later ranked first in WorldArena 2.0. Yingsu also launched an open data initiative.

Large Language Models · Research · Computer Vision

Sun Yat-sen Researchers Propose SpatialSV to Strengthen Multimodal LLMs’ Spatial Intelligence

A Sun Yat-sen team's SpatialSV embeds explicit 3D geometry in MLLMs, improving spatial accuracy and halving manual annotation.

AI Chips & Compute · Products & Applications · Computer Vision

New AI Chip and 3D Scene Model Target Generation Bottlenecks

AI video chip SmarCo GC3 and 3D scene model WorldGen launch to ease content creation bottlenecks.

Products & Applications · Computer Vision

Dyson Unveils $499 CameraJet Smart Toothbrush With AI-Powered Mouthwash Flossing

Dyson unveils $499 CameraJet, a smart toothbrush that uses a camera and on-device AI to detect gaps between teeth and automatically blast mouthwash into them.

Products & Applications · Computer Vision

Ex-DJI Engineers Launch 'World's First Robotic Extreme Zoom Camera' on Kickstarter

FarseerTech, founded by ex-DJI engineers, is crowdfunding the RocXZoom, a robotic extreme zoom camera with 50x hybrid zoom and AI bird recognition.

Computer Vision · Companies

Nvidia's DLSS 5 launches with NBA 2K27, but only for RTX 50-series

Nvidia's DLSS 5 debuts with NBA 2K27 on Sep 3, RTX 50-series only; mixed early reviews cite realism and performance hit.

Products & Applications · Computer Vision · Companies

Dyson unveils CameraJet toothbrush that uses a camera and jet spray to floss while you brush

Dyson has launched the CameraJet, a camera-equipped electric toothbrush that uses AI-guided mouthwash jets to floss between teeth while brushing. The premium device, priced around $500 in the U.S., went on sale Tuesday.

Research · Robotics · Computer Vision

TUM Professor and New TrAct Paper Target AI's Deployment Gap

Leiphone reports: TUM's Ziyue Li and a new Li Fei-Fei/Wu Jiajun paper both tackle AI's gap between benchmark success and real-world deployment.

Products & Applications · Computer Vision

Free open-source app Paperwork helps organize and search documents on Linux and Windows

A free, open-source tool called Paperwork helps Linux and Windows users tag, label, and search documents locally, according to ZDNet.

Computer Vision · Companies

Amap Unveils ABot-Recon for Real-Time 3D Reconstruction of 10,000-Frame Scenes

Amap, an Alibaba Group subsidiary, unveiled ABot-Recon, a streaming 3D reconstruction model that rebuilds 10,000-frame scenes in real time using just 12 local frames and no long-range memory.

Large Language Models · Computer Vision

OpenAI reboot, Salesforce-Anthropic 'Claudeforce' and Alibaba Qoder mark a busy Aug. 27

OpenAI disclosed strategic missteps and a safety-focused reboot, Salesforce and Anthropic unveiled 'Claudeforce,' Alibaba launched Qoder, and Chinese firms pushed AI edge and industry applications.

Products & Applications · Robotics · Computer Vision

Ex-Meta scientists launch open-weight visual AI model Isaac 0.5 for industrial robots

Perceptron, founded by former Meta researchers, launched Isaac 0.5, an open-weight vision model to help robots perceive, reason and act in warehouses and factories.

Companies · Computer Vision

OBSBOT co-founder: Imaging automation is the endgame as company leads high-end webcams

OBSBOT, a Shenzhen AI camera maker, has become the global leader in high-end webcams with over 50% market share, according to co-founder Li Liang. The company sustained 50% annual growth and recently completed a Series D round.

Computer Vision · Research

LiveEdit: Real-Time Streaming Video Editing Framework Accepted at ECCV 2026

Tsinghua and HKUST researchers propose LiveEdit, a streaming video editing framework achieving 12.66 FPS with 4-step denoising, enabling real-time editing for live scenarios. Accepted at ECCV 2026.

Robotics · Computer Vision

Researchers devise sonar-guided system to help underwater vehicles see through murky water

Underwater vehicles could navigate murky water using a sonar-guided technique from WHOI researchers, MIT Tech Review reports.

Computer Vision

AI Won't Replace Radiologists, But It Will Dramatically Change Their Jobs

AI is not replacing radiologists, but it is changing their roles. Three-quarters of FDA-cleared AI medical devices target radiology, and AI tools are improving efficiency and accuracy in image interpretation.

Products & Applications · Large Language Models · Computer Vision

Google Search introduces five AI-powered tools for home decor

Google Search has rolled out new features, including AI Mode, Lens, Circle to Search, Search Live, and price tracking, to help users with home decor projects as related searches surge.

Products & Applications · Computer Vision

Small Screens, Tiny Cameras: How AI Hardware Is Learning to See Less but Understand More

Small e-ink displays and Apple's 1-megapixel AirPods camera show AI hardware is going low-res and ambient.

Products & Applications · Computer Vision

deepDoctection 1.2.x Tutorial Shows End-to-End Document Intelligence Pipeline

MarkTechPost details a deepDoctection 1.2.x workflow that combines layout detection, table recognition, OCR, reading-order reconstruction and JSONL export for RAG systems.

Companies · Large Language Models · Computer Vision

Ordnance Survey CEO: Location is the obvious way to connect data in AI era

Ordnance Survey CEO Nick Bolton discusses transforming the 225-year-old institution into an AI powerhouse, emphasizing location data as a universal connector.

Products & Applications · Large Language Models · Computer Vision

AI-Powered Fitness Coach BodyPark Atom Impresses Reviewer with Movement Mapping

A TechRadar journalist tested the BodyPark Atom, an AI-powered home fitness coach, and was impressed by its movement mapping despite initial skepticism.

Large Language Models · Computer Vision

DeepSeek launches V4 Flash Vision Exp multimodal model, beating Opus 4.8 on visual tests

DeepSeek has introduced V4 Flash Vision Exp, a multimodal LLM that outperformed Anthropic's Opus 4.8 on two image benchmarks, with a paid platform launch and possible open-source release later.

Large Language Models · Computer Vision

SenseTime Open-Sources SenseNova U1.5 Lite with Ultra-Long Instruction Support and Native 4K

SenseTime has open-sourced SenseNova U1.5 Lite, a lightweight multimodal model supporting 3-4k character instructions and native 4K output for stable visual creation.

Large Language Models · Computer Vision

DeepSeek Releases Multimodal Model deepseek-v4-flash-vision-exp, Adds Harness Support

DeepSeek unveils experimental vision model deepseek-v4-flash-vision-exp for agents, with Harness support and Files API.