↓ Skip to main content

Nano Banana 2.1 (gemini-nano-banana-2.1) Model Card #

Overview #

Nano Banana 2.1 (gemini-nano-banana-2.1) is Google’s next-generation multimodal creative workhorse built on the Gemini 3.6 Flash architecture. Released on October 6, 2026, it delivers enhanced spatial composition, seamless panoramic synthesis across 14 aspect ratios, mask-based conversational editing, three-tier controllable thinking (minimal, medium, high), real-time Google Web and Image Search grounding, and multi-reference character consistency across up to 14 anchors.

  • Model Code: gemini-nano-banana-2.1
  • Architecture: Gemini 3.6 Flash
  • Release Date: October 6, 2026
  • Status: General Availability (GA)
  • CLI Aliases: nano-banana-2.1 (default), nano-banana-2-1

Technical Specifications #

PropertyValue
Input ModalitiesText, Images (up to 14), Video (YouTube URLs, MP4s), PDF
Output ModalitiesImage and Text (Interleaved)
Input Token Limit131,072 tokens
Output Token Limit32,768 tokens
Supported Resolutions512px (0.5K), 1K (1024px), 2K (2048px), 4K (4096px)
Aspect RatiosAll 14 discrete ratios (1:1, 1:4, 1:8, 2:3, 3:2, 3:4, 4:1, 4:3, 4:5, 5:4, 8:1, 9:16, 16:9, 21:9) with panoramic seam/tiling elimination
Thinking ProcessSupported (minimal [default], medium, high)
Search GroundingGoogle Web Search (--search) + Google Image Search (--image-search)
Reference AnchorsUp to 14 reference images (up to 10 object anchors + up to 4 character resemblance anchors)
Conversational EditingMulti-turn mask-based inpainting, outpainting, and local regional edits
WatermarkingSynthID (Always On) + C2PA Metadata

Complete Resolution Grid & Token Consumption #

Aspect Ratio512px (0.5K)Tokens1K ResolutionTokens2K ResolutionTokens4K ResolutionTokens
1:1512 × 5127471024 × 10241,1202048 × 20481,6804096 × 40962,520
1:4256 × 1024747512 × 20481,1201024 × 40961,6802048 × 81922,520
1:8192 × 1536747384 × 30721,120768 × 61441,6801536 × 122882,520
2:3424 × 632747848 × 12641,1201696 × 25281,6803392 × 50562,520
3:2632 × 4247471264 × 8481,1202528 × 16961,6805056 × 33922,520
3:4448 × 600747896 × 12001,1201792 × 24001,6803584 × 48002,520
4:11024 × 2567472048 × 5121,1204096 × 10241,6808192 × 20482,520
4:3600 × 4487471200 × 8961,1202400 × 17921,1204800 × 35842,520
4:5464 × 576747928 × 11521,1201856 × 23041,6803712 × 46082,520
5:4576 × 4647471152 × 9281,1202304 × 18561,6804608 × 37122,520
8:11536 × 1927473072 × 3841,1206144 × 7681,68012288 × 15362,520
9:16384 × 688747768 × 13761,1201536 × 27521,6803072 × 55042,520
16:9688 × 3847471376 × 7681,1202752 × 15361,6805504 × 30722,520
21:9792 × 1687471584 × 6721,1203168 × 13441,6806336 × 26882,520

Capabilities & Key Features #

1. Three-Tier Controllable Thinking (minimal, medium, high) #

Nano Banana 2.1 introduces the medium thinking level on top of minimal and high:

  • "minimal": Fast synthesis with minimal pre-generation reasoning latency.
  • "medium": Balanced spatial layout planning, typography verification, and perspective coherence with sub-second reasoning overhead (exclusive to gemini-nano-banana-2.1).
  • "high": Deep compositional planning for multi-character scenes, intricate architectural perspectives, complex lighting, and dense infographic text.

2. Panoramic Seam & Tiling Elimination (1:4, 4:1, 1:8, 8:1) #

Built on Gemini 3.6 Flash attention windows, Nano Banana 2.1 eliminates repetition seams and horizon discontinuities across extreme panoramic (4:1, 8:1) and vertical scroll (1:4, 1:8) ratios, making it ideal for 2D game parallax backgrounds, seamless level strips, and website hero banners.

3. Mask-Based Conversational Editing & Multi-Reference Consistency #

Pass up to 14 reference images (up to 10 object references + 4 character anchors) alongside optional binary/alpha mask inputs to perform surgical regional inpainting while preserving surrounding pixels and character identity.

4. Google Web Search & Image Search Grounding #

Query Google Web Search and Google Image Search in real time to ground generated visuals in verified facts, real-world landmarks, current events, and accurate visual taxonomy.

5. Video Context Understanding #

Pass YouTube URLs or MP4 files to synthesize keyframe illustrations, sprite sheets, or promotional posters grounded in temporal video context.


SDK Code Examples #

python
import base64
from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input="A seamless 2D side-scrolling pixel-art parallax background of a bioluminescent crystal cavern",
    tools=[{"type": "google_search", "search_types": ["web_search", "image_search"]}],
    generation_config={"thinking_level": "medium"},
    response_format={
        "type": "image",
        "aspect_ratio": "4:1",
        "image_size": "4K",
    },
)

if interaction.output_image:
    with open("cavern_parallax_4k.png", "wb") as f:
        f.write(base64.b64decode(interaction.output_image.data))

2. Python SDK (google-genai) — Multi-Reference Character Consistency & Mask-Based Editing #

python
import base64
from google import genai

client = genai.Client()

def encode_png(path: str) -> str:
    with open(path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input=[
        {"type": "image", "data": encode_png("references/chibi-dani.png"), "mime_type": "image/png"},
        {"type": "image", "data": encode_png("references/celebrating-dani.png"), "mime_type": "image/png"},
        {"type": "image", "data": encode_png("references/speaker-dani.png"), "mime_type": "image/png"},
        {
            "type": "text",
            "text": (
                "Keep the exact character identity, purple hair, glasses, and chibi proportions from the reference images. "
                "Generate a 4-frame horizontal sprite sheet of Dani casting a glowing lightning spell, solid magenta (#FF00FF) background."
            ),
        },
    ],
    generation_config={"thinking_level": "high"},
    response_format={"type": "image", "aspect_ratio": "4:1", "image_size": "2K"},
)

if interaction.output_image:
    with open("dani_spell_sheet.png", "wb") as f:
        f.write(base64.b64decode(interaction.output_image.data))

3. JavaScript / TypeScript SDK (@google/genai) #

typescript
import { GoogleGenAI } from "@google/genai";
import * as fs from "node:fs";

const ai = new GoogleGenAI({});

async function generateBanner() {
  const interaction = await ai.interactions.create({
    model: "gemini-nano-banana-2.1",
    input: "A vibrant isometric cyber-arcade cabinet with glowing neon marquee",
    generation_config: {
      thinking_level: "medium",
    },
    response_format: {
      type: "image",
      aspect_ratio: "16:9",
      image_size: "2K",
    },
  });

  if (interaction.output_image) {
    fs.writeFileSync(
      "arcade_cabinet.png",
      Buffer.from(interaction.output_image.data, "base64")
    );
  }
}

generateBanner();

4. Go SDK (google.golang.org/genai) #

go
package main

import (
	"context"
	"fmt"
	"log"
	"os"

	"google.golang.org/genai"
)

func main() {
	ctx := context.Background()
	client, err := genai.NewClient(ctx, nil)
	if err != nil {
		log.Fatalf("failed to create genai client: %v", err)
	}

	resp, err := client.Models.GenerateContent(
		ctx,
		"gemini-nano-banana-2.1",
		genai.Text("A clean 16-bit SNES style treasure chest sprite on a solid #FF00FF background"),
		&genai.GenerateContentConfig{
			ResponseModalities: []string{"IMAGE"},
		},
	)
	if err != nil {
		log.Fatalf("generation failed: %v", err)
	}

	for _, part := range resp.Candidates[0].Content.Parts {
		if part.InlineData != nil {
			if err := os.WriteFile("chest_sprite.png", part.InlineData.Data, 0644); err != nil {
				log.Fatalf("write file failed: %v", err)
			}
			fmt.Println("Saved chest_sprite.png")
			break
		}
	}
}