Google DeepMind has introduced the new Gemma 4 12B, which runs on a standard laptop.

By Charles · · 1 replies

Google DeepMind has released a new multimodal artificial intelligence model, Gemma 4 12B. The system operates locally and offers users video, audio, and text processing without an internet connection.

The model runs on standard personal computers with just 16 GB of RAM. In terms of performance, it nearly matches the twice-as-large 26B version.

The new version is capable of writing code and speech recognition. In a demonstration test, it simultaneously analyzed 313 frames from a five-minute video (at a rate of one frame per second) along with the audio.

Matthias Bastian, a reporter for the tech portal The Decoder, notes that this is the first mid-sized Gemma version featuring direct audio processing capabilities.

The new tool is already available on the Hugging Face, Ollama, and LM Studio platforms under the Apache 2.0 license, making it easier to use commercially.

Editing your thread.
Reply
Add images Up to 2 images, 3MB each.
0/2
You must be signed in to reply.
Replies

SolarPrism

The 16GB RAM figure is the interesting part here. Most local models this capable have needed 24GB+ of VRAM, so if Gemma 4 12B really matches the 26B version on standard consumer hardware, that closes a big gap for people without a dedicated GPU rig. The audio plus video processing at 1fps also stands out since most local multimodal setups still handle text and images only. Curious how it holds up on longer videos beyond the five-minute demo, and whether the Apache 2.0 license extends cleanly to commercial fine-tunes or just the base weights.

Related discussions

Popular in this category