Skip to main content

Multimodal Capabilities Guide

NeuroLink provides comprehensive multimodal support, allowing you to combine text with various media types in a single AI interaction. This guide covers all supported input types, provider capabilities, and best practices.

Overview​

Supported Input Types:

  • Images - JPEG, PNG, GIF, WebP, AVIF, HEIC (vision-capable models)
  • PDFs - Document analysis and content extraction
  • CSV/Spreadsheets - Data analysis and tabular content processing
  • Audio - Transcription, analysis, and real-time voice input (Audio Input Guide)
  • Documents - Excel, Word, RTF, OpenDocument formats (File Processors Guide)
  • Data Files - JSON, YAML, XML with validation and formatting
  • Markup - HTML, SVG, Markdown with security sanitization
  • Source Code - 50+ programming languages with syntax detection

All multimodal inputs work seamlessly across both the CLI and SDK, with automatic format detection and provider-specific optimization.

New in 2026: NeuroLink now supports 17+ file types through the ProcessorRegistry system. See the File Processors Guide for comprehensive documentation.


Provider Support Matrix​

Not all providers support all multimodal capabilities. Use this matrix to select the right provider for your use case.

Vision (Images)​

ProviderSupportedRecommended ModelsMax ImagesMax SizeNotes
OpenAI✅gpt-4o, gpt-4o-mini, gpt-5.210~20 MBBest for general vision tasks
Azure OpenAI✅gpt-4o, gpt-4o-mini10~20 MBSame as OpenAI
Google AI Studio✅gemini-2.5-pro, gemini-2.5-flash, gemini-3-flash16~20 MBExcellent for visual reasoning
Google Vertex AI✅gemini-2.5-pro, gemini-2.5-flash, Claude models16/20~20 MBGemini: 16 images, Claude: 20 images
Anthropic✅claude-3.5-sonnet, claude-3.7-sonnet20~20 MBStrong visual understanding
AWS Bedrock✅Claude models20~20 MBSame as Anthropic
Ollama✅llava, bakllava, llava-phi310VariesLocal vision models
LiteLLM✅Depends on upstream10VariesProxy to vision-capable models
Mistral✅pixtral-12b-2409, pixtral-large-241110~20 MBMultimodal Mistral models
OpenRouter✅Depends on model10VariesRoutes to various vision models
Hugging Face⚠️LimitedVariesVariesModel-dependent
AWS SageMaker❌N/A--Not supported
OpenAI Compatible⚠️Depends on endpointVariesVariesServer-dependent

Legend:

  • ✅ Full support with multiple models
  • ⚠️ Limited or server-dependent support
  • ❌ Not supported

PDF Documents​

ProviderSupportedMax SizeMax PagesProcessing ModeNotes
Google Vertex AI✅5 MB100Native PDFBest for document analysis
Anthropic✅5 MB100Native PDFClaude excels at document understanding
AWS Bedrock✅5 MB100Native PDFVia Claude models
Google AI Studio✅2000 MB100Native PDFHandles very large files
OpenAI✅10 MB100Files APIgpt-4o, gpt-4o-mini, o1
Azure OpenAI✅10 MB100Files APIUses OpenAI Files API
LiteLLM✅10 MB100ProxyDepends on upstream model
OpenAI Compatible✅10 MB100VariesServer-dependent
Mistral✅10 MB100Native PDFNative support
Hugging Face✅10 MB100Model-dependentVaries by model
Ollama❌---Not supported
OpenRouter⚠️VariesVariesDepends on modelRoute-dependent
AWS SageMaker❌---Not supported

CSV/Spreadsheet Data​

ProviderSupportedMax RowsFormat OptionsNotes
Supported Providers✅10,000raw, json, markdownBroad support - processed as text

CSV support works across supported providers because files are converted to text before sending to the AI model. The file is parsed and formatted (raw CSV, JSON, or Markdown table) before inclusion in the prompt.

Format Recommendations:

  • Raw format - Best for large files (minimal token usage)
  • JSON format - Best for structured data processing
  • Markdown format - Best for small datasets (<100 rows), readable tables

Audio Input​

ProviderNative AudioTranscriptionReal-timeMax DurationNotes
Google AI Studio✅✅✅1 hourBest for real-time voice
Google Vertex AI✅✅✅1 hourNative Gemini audio support
OpenAI❌✅ Whisper❌25 MBExcellent transcription accuracy
Azure OpenAI❌✅ Whisper❌25 MBVia Whisper integration
Anthropic❌Via fallback❌-Uses transcription approach
AWS Bedrock❌Via fallback❌-Uses transcription approach
Others❌Via fallback❌-Audio transcribed before processing

For comprehensive audio documentation, see the Audio Input Guide.


Video​

Two mechanisms, and which one a provider gets changes what it can answer.

Native (inline) — the video file itself travels in the request. The model sees continuous motion and hears the audio track. Only Google's Gemini front ends accept this today.

Frame extraction — ffmpeg pulls keyframes at intervals and they are attached as ordinary images, alongside a metadata summary in the prompt text. The model sees a handful of stills and hears nothing.

ProviderMechanismInline ceilingAudio trackDefault frame budget
Google AI StudioNative (inline)14 MB✅ heard16 (fallback only)
Google Vertex AI (Gemini models)Native (inline)14 MB✅ heard16 (fallback only)
OpenAI / Azure OpenAIFrame extraction–❌8
Anthropic / AWS BedrockFrame extraction–❌8
MistralFrame extraction–❌8
LiteLLM / OpenRouterFrame extraction–❌8
Ollama / llama.cppFrame extraction–❌4

The 14 MB ceiling is one shared source-byte budget for the whole request, not a separate per-clip limit. Before a clip ships natively, its bytes are added to the source bytes of every inline part already in the request, and the clip is delivered only if the total stays within 14 MB. Those parts include images (keyframes among them), PDFs, audio and any clip accepted earlier. With nothing else attached, that reduces to the clip's own size. The budget is in source bytes because inline data travels as base64 in a JSON body, which costs four bytes per three, and Gemini rejects a request over 20 MB. 14 MB of source encodes to ~18.7 MB, leaving headroom for the prompt and the request envelope. 15 MB would encode to exactly 20.0 MB, with nothing left over. A PDF that ends up sent as extracted text may still be counted, which errs toward an unnecessary keyframe fallback rather than an oversized request. Falling back is not an error: the clip goes as keyframes, and a line naming the reason is logged.

Whichever mechanism applies, the metadata summary (duration, resolution, codec, frame rate) is always folded into the prompt text.

ffmpeg is optional, and its absence is not symmetric. Frame extraction shells out to ffmpeg at runtime. Without it a video attached to a frame-extraction provider yields the metadata summary and nothing visual — no error, just an empty keyframes array. Native delivery needs no ffmpeg at all, so on a machine without it Gemini still receives the whole clip.

Supported containers​

.mp4, .webm, .mov, .avi, .mkv, .mpeg, .flv, .wmv, .3gp and the other formats in the detector's video registry are all accepted for frame extraction. Inline delivery is narrower — it is limited to the containers Gemini reads directly (mp4, mpeg, mov, avi, flv, webm, wmv, 3gpp) — and a container outside that set falls back to keyframes rather than being re-encoded. Transcoding a long recording mid-request costs minutes of CPU for a payload that would usually breach the inline ceiling anyway.

Image Input​

Quick Start​

CLI:

# Single image
npx @juspay/neurolink generate "Describe this interface" \
--image ./designs/dashboard.png --provider google-ai

# Remote URL
npx @juspay/neurolink generate "Analyze this diagram" \
--image https://example.com/architecture.png --provider openai

# Multiple images
npx @juspay/neurolink generate "Compare these screenshots" \
--image ./before.png \
--image ./after.png \
--provider anthropic

SDK:

import { readFileSync } from "node:fs";
import { NeuroLink } from "@juspay/neurolink";

const neurolink = new NeuroLink({ enableOrchestration: true });

const result = await neurolink.generate({
input: {
text: "Analyze these product screenshots",
images: [
readFileSync("./homepage.png"), // Local file as Buffer
"https://example.com/chart.png", // Remote URL
],
},
provider: "google-ai",
});

Image Formats Supported​

Accepted formats:

  • JPEG (.jpg, .jpeg)
  • PNG (.png)
  • GIF (.gif)
  • WebP (.webp)
  • AVIF (.avif) - detected from content (avif/avis/avio brands) as well as extension
  • BMP (.bmp), TIFF (.tif, .tiff) - detected from content
  • HEIC (.heic, .heif) - detected from HEIC/HEIF content brands as well as extension; unsupported provider formats still require PNG/JPEG conversion

The MIME type is sniffed from the buffer's magic bytes, not assumed from the filename. A buffer whose bytes match no known image format is labeled application/octet-stream (with a warning) rather than silently mislabeled as JPEG.

Input methods:

  • Buffer objects - readFileSync() from Node.js
  • Local file paths - Relative or absolute paths
  • HTTPS URLs - Remote images (auto-downloaded)

Image Alt Text (Accessibility)​

NeuroLink supports alt text for images, improving accessibility and providing additional context to AI models.

const result = await neurolink.generate({
input: {
text: "Compare these revenue charts",
images: [
{
data: readFileSync("./q1-revenue.png"),
altText: "Q1 2024 revenue chart showing 15% growth",
},
{
data: "https://example.com/q2-revenue.png",
altText: "Q2 2024 revenue chart showing 22% growth",
},
],
},
provider: "openai",
});

Alt text best practices:

  • Keep concise (under 125 characters ideal)
  • Focus on key information the image conveys
  • Alt text is automatically included as context in prompts

Image Size Limits​

Provider-specific limits:

  • Most providers: ~20 MB per image
  • Recommended: Resize images to < 2 MP for faster processing
  • Token usage: ~7,000 tokens per image (varies by provider)

Optimization tips:

  • Compress images before sending for large batches
  • Use appropriate resolution (1920x1080 often sufficient)
  • Pre-process images to reduce unnecessary detail

PDF Document Input​

Quick Start​

CLI:

# Auto-detect PDF
npx @juspay/neurolink generate "Summarize this report" \
--file ./financial-report.pdf --provider vertex

# Explicit PDF
npx @juspay/neurolink generate "Extract key terms from contract" \
--pdf ./contract.pdf --provider anthropic

# Multiple PDFs
npx @juspay/neurolink generate "Compare these documents" \
--pdf ./version1.pdf \
--pdf ./version2.pdf \
--provider vertex

SDK:

// Auto-detect (recommended)
await neurolink.generate({
input: {
text: "Analyze this document",
files: ["./report.pdf", "./data.csv"], // Mixed file types
},
provider: "vertex",
});

// Explicit PDF
await neurolink.generate({
input: {
text: "Compare Q1 and Q2 reports",
pdfFiles: ["./q1-report.pdf", "./q2-report.pdf"],
},
provider: "anthropic",
});

PDF Processing Modes​

Provider-specific approaches:

ProviderModeToken UsageBest For
Vertex AI, Anthropic, BedrockNative PDF~1,000 tokens/3 pagesVisual + text extraction
Google AI StudioNative PDF~1,000 tokens/3 pagesLarge files (up to 2 GB)
OpenAI, AzureFiles API~1,000 tokens/3 pagesText-only mode optimal

Visual vs. Text-only mode:

  • Visual mode: Preserves layout, tables, charts (~7,000 tokens/3 pages)
  • Text-only mode: Extracts text content only (~1,000 tokens/3 pages)

PDF Best Practices​

  • Choose the right provider: Vertex AI or Anthropic for best results
  • Check file size: Most providers limit to 5 MB (AI Studio supports 2 GB)
  • Use streaming: For large documents, streaming provides faster initial results
  • Combine with other files: Mix PDFs with CSV data and images
  • Be specific in prompts: "Extract all monetary values" vs. "Tell me about this PDF"
  • Set appropriate token limits: Recommended 2000-8000 tokens for PDF analysis

CSV/Spreadsheet Input​

Quick Start​

CLI:

# Auto-detect CSV
npx @juspay/neurolink generate "Analyze sales trends" \
--file ./sales_2024.csv

# Explicit CSV with options
npx @juspay/neurolink generate "Summarize data" \
--csv ./data.csv \
--csv-max-rows 500 \
--csv-format raw

SDK:

// Auto-detect (recommended)
await neurolink.generate({
input: {
text: "Analyze this sales data",
files: ["./sales.csv"], // Auto-detected as CSV
},
});

// Explicit CSV with options
await neurolink.generate({
input: {
text: "Compare quarterly data",
csvFiles: ["./q1.csv", "./q2.csv"],
},
csvOptions: {
maxRows: 1000,
formatStyle: "json", // or "raw", "markdown"
},
});

CSV Format Options​

Three format styles:

  1. Raw format (default)

    • Best for large files
    • Minimal token usage
    • Preserves original CSV structure
    name,age,city
    Alice,30,NYC
    Bob,25,LA
  2. JSON format

    • Structured data processing
    • Easier for AI to parse
    • Higher token usage
    [
    { "name": "Alice", "age": 30, "city": "NYC" },
    { "name": "Bob", "age": 25, "city": "LA" }
    ]
  3. Markdown format

    • Readable tables
    • Good for small datasets (<100 rows)
    • Moderate token usage
    | name  | age | city |
    | ----- | --- | ---- |
    | Alice | 30 | NYC |
    | Bob | 25 | LA |

CSV Configuration​

const result = await neurolink.generate({
input: {
text: "Analyze customer data",
csvFiles: ["./customers.csv"],
},
csvOptions: {
maxRows: 1000, // Limit rows (default: 1000, max: 10000)
formatStyle: "json", // Format: "raw" | "json" | "markdown"
includeHeaders: true, // Include header row (default: true)
},
});

CSV Best Practices​

  • Use raw format for large files to minimize token usage
  • Use JSON format for structured processing when AI needs to manipulate data
  • Limit to 1000 rows by default (configurable up to 10,000)
  • Combine CSV with visualization images for comprehensive analysis
  • Works across supported providers (not just vision-capable models)

Video Input​

SDK usage​

// Auto-detect: the same `files` array used for every other format
const result = await neurolink.generate({
input: {
text: "What happens in this clip, and what is said?",
files: ["./demo.mp4"],
},
provider: "google-ai",
});

// Explicit: `videoFiles` is folded into `files` before detection runs
const result = await neurolink.generate({
input: {
text: "Summarise this recording",
videoFiles: ["./standup.mp4"],
},
provider: "vertex",
});

Frame-extraction options​

These control the keyframes and are ignored by a provider on the native path.

const result = await neurolink.generate({
input: { text: "Describe each scene", files: ["./demo.mp4"] },
provider: "openai",
videoOptions: {
frames: 16, // Keyframe budget. Clamped to the processor ceiling of 100.
quality: 90, // Encoder quality 1-100. Default 80.
format: "jpeg", // "jpeg" | "png". Default jpeg.
// Transcribe the clip's speech with OpenAI Whisper and include it as
// text. Needs ffmpeg and OPENAI_API_KEY; see below.
transcribeAudio: true,
},
});

Without an explicit frames, the interval is chosen from the clip's duration — roughly every second for clips under 10s, widening to every three minutes for recordings over half an hour, always capped at 100 frames.

CLI usage​

# Auto-detect
neurolink generate "Describe this video" --file demo.mp4 --provider google-ai

# Explicit flag
neurolink generate "Summarise this" --video standup.mp4 --provider vertex

# Frame-extraction knobs
neurolink generate "Describe each scene" --file demo.mp4 \
--provider openai --video-frames 16 --video-quality 90 --video-format jpeg

--transcribe-audio (or videoOptions.transcribeAudio) gives a frame-path provider the clip's speech, which keyframes cannot carry. The audio track is extracted with ffmpeg, transcribed by OpenAI Whisper (whisper-1, using OPENAI_API_KEY, and OPENAI_BASE_URL when set), and sent alongside the keyframes under a Spoken Audio (transcribed) heading. Transcription is additive: when it cannot run — the clip has no audio track, OPENAI_API_KEY is unset, the extracted audio is over Whisper's 25 MB limit, or extraction or the upload fails — the request still succeeds with the keyframes and metadata, and a No transcript for <file>: <reason> warning says which. A Gemini provider does not need it while the clip is delivered natively, since that includes the audio track; a clip that falls back to keyframes (over the inline ceiling) loses its speech unless --transcribe-audio is set.

Embedded subtitle tracks are extracted separately and included whenever they exist, with or without transcribeAudio.

Asking which mechanism applies​

The capability table is exported, so a caller can ask before sending — useful for choosing a provider, or for budgeting.

import {
estimateVideoTokens,
getVideoProviderConfig,
supportsNativeVideo,
} from "@juspay/neurolink";

supportsNativeVideo("google-ai"); // true
supportsNativeVideo("openai"); // false
getVideoProviderConfig("nope"); // null — not described, not "takes frames"

// Duration drives the native price; frame count drives the other one.
estimateVideoTokens({ provider: "google-ai", durationSec: 60 }); // ~15,600
estimateVideoTokens({ provider: "openai", durationSec: 60, frameCount: 8 }); // ~2,000

Best practices​

  • Ask an audio question of a Gemini provider only. Nothing on the frame path can hear the clip, so "what was said" has no answer there.
  • Keep clips under 14 MB when you want native delivery. Re-encoding a screen recording at a lower bitrate usually gets a long clip under the ceiling without losing anything a model would read.
  • Raise frames rather than quality for detail over time. Frames cost roughly 250 tokens each regardless of quality; more of them is what buys coverage of a clip where things change.
  • Lower frames for local runtimes. Ollama and llama.cpp hold every frame in the same memory as the model; four is the practical default.

Troubleshooting​

SymptomCauseFix
Model describes the file but not its contentNo keyframes extracted — ffmpeg missing — and the provider takes framesInstall ffmpeg, or use a Gemini provider, which needs none
Sending keyframes instead of the clip in the logsThe clip failed the inline gate; the reason is on the same lineShorten or re-encode it, or accept the frames
Model cannot hear speech on a Gemini providerThe clip exceeded the inline ceiling and fell back to framesGet the source under 14 MB, or pass --transcribe-audio
--transcribe-audio produced no transcriptThe No transcript for warning names the reason: no audio track, OPENAI_API_KEY unset, or audio over Whisper's 25 MB limitSet OPENAI_API_KEY, or split the recording
Only the first seconds of a long clip are describedThe frame budget was hit before the endRaise frames; the interval then spreads evenly across the whole clip

Combining Multiple Input Types​

NeuroLink excels at combining different media types in a single request.

Mixed Media Example​

const result = await neurolink.generate({
input: {
text: "Analyze this product launch: review the presentation, compare sales data, and assess the promotional materials",
pdfFiles: ["./presentation.pdf"], // Slides
csvFiles: ["./sales-data.csv"], // Numbers
images: [
readFileSync("./promo-banner.png"), // Marketing material
"https://example.com/ad-campaign.jpg",
],
},
provider: "vertex", // Supports all input types
});

Streaming with Multimodal​

const stream = await neurolink.stream({
input: {
text: "Analyze this floor plan and cost breakdown",
images: ["./floor-plan.jpg"],
csvFiles: ["./costs.csv"],
},
provider: "google-ai",
});

for await (const chunk of stream) {
process.stdout.write(chunk.text ?? "");
}

Batch with Multimodal (CLI)​

The batch command supports --image, --csv, --pdf, and --video. The file(s) are attached identically to every prompt in the batch (a one-line notice is printed to stderr, unconditionally — it is not suppressed by --quiet, so it never corrupts --format json output written to stdout):

neurolink batch questions.txt --csv sales.csv --format json

--file (auto-detect) is not available in batch, because it collides with the <file> positional (the prompts-list path). Use the explicit --image / --csv / --pdf / --video flags instead.

File validation & troubleshooting (CLI)​

Before any provider call, the CLI validates local file inputs across generate, stream, and batch:

  • A path that points at a directory, doesn't exist, or can't be read (e.g. a permissions error) is rejected up front with a clear error and troubleshooting hints — no cryptic EISDIR/EACCES deep in processing. This also applies to the batch <file> prompts-list positional itself.
  • A large file (images > 10 MB, CSV/PDF > 50 MB) prints a non-blocking warning to stderr. Like the batch attachment notice, this is unconditional — visible without --debug and regardless of --quiet — and never mixes into stdout, so --format json output stays valid JSON even when large-file warnings fire.

This runs even under --dry-run and without API keys configured.


Configuration & Fine-tuning​

Image-Specific Options​

const result = await neurolink.generate({
input: {
text: "Analyze these screenshots",
images: [
{
data: readFileSync("./screenshot.png"),
altText: "Product dashboard showing KPIs",
},
],
},
provider: "openai",
maxTokens: 2000, // Increase for detailed image analysis
});

PDF-Specific Options​

const result = await neurolink.generate({
input: {
text: "Extract financial data from this report",
pdfFiles: ["./annual-report.pdf"],
},
provider: "vertex",
maxTokens: 8000, // Large token budget for comprehensive extraction
});

Regional Routing​

Some providers require regional configuration for optimal performance:

const result = await neurolink.generate({
input: {
text: "Analyze this document",
pdfFiles: ["./contract.pdf"],
},
provider: "vertex",
region: "us-central1", // Vertex AI region
});

Best Practices​

General Guidelines​

  1. Provide descriptive prompts - Reference specific images/files by name
  2. Use alt text for accessibility - Helps both AI and screen readers
  3. Combine analytics + evaluation - Benchmark multimodal quality before production
  4. Cache remote assets locally - Avoid repeated downloads for frequently used files
  5. Stream for user-facing apps - Use generate() for structured JSON output

Image Best Practices​

  • Provide short captions describing each image in the prompt
  • Pre-compress large images to reduce processing time
  • Use appropriate image formats (JPEG for photos, PNG for diagrams)
  • Consider token limits when sending multiple images

PDF Best Practices​

  • Choose providers with native PDF support (Vertex, Anthropic, Bedrock)
  • Be specific about what you need extracted
  • Use streaming for large documents
  • Set appropriate maxTokens (2000-8000 recommended)

CSV Best Practices​

  • Use raw format for large datasets
  • Use JSON format when AI needs structured data manipulation
  • Limit rows to avoid token exhaustion
  • Combine with images for visual + numerical analysis

Troubleshooting​

Common Issues​

IssueSolution
"Image not found"Check file paths are relative to CWD where CLI is invoked
"Provider does not support images"Switch to vision-capable provider (see matrix above)
"Error downloading image"Ensure URL returns HTTP 200 and doesn't require authentication
"Large response latency"Pre-compress images and reduce resolution to < 2 MP
"Streaming ends early"Disable tools (--disableTools) to avoid tool call interruptions
"PDF too large"Use Google AI Studio (2 GB limit) or split into smaller chunks
"CSV token overflow"Reduce maxRows or use raw format instead of JSON/markdown

Provider-Specific Issues​

OpenAI/Azure:

  • Images must be < 20 MB
  • PDFs processed via Files API (may take longer)

Google AI Studio/Vertex:

  • Best for large PDFs (AI Studio supports up to 2 GB)
  • Gemini models have excellent visual reasoning

Anthropic/Bedrock:

  • Claude excels at document understanding
  • Strong visual and text analysis capabilities

Ollama:

  • Use vision-capable models like llava, bakllava
  • Local processing - no cloud API required

Document Processing:

Q4 2025 Features:

Advanced Features:

Documentation:


Examples & Recipes​

Example 1: Product Analysis​

Analyze a product page with screenshot, description, and pricing data:

const analysis = await neurolink.generate({
input: {
text: "Analyze this product: review the screenshot, pricing data, and provide recommendations",
images: [readFileSync("./product-screenshot.png")],
csvFiles: ["./pricing-tiers.csv"],
},
provider: "google-ai",
maxTokens: 3000,
});

Example 2: Document Comparison​

Compare two versions of a contract:

const comparison = await neurolink.generate({
input: {
text: "Compare these two contract versions and highlight key differences",
pdfFiles: ["./contract-v1.pdf", "./contract-v2.pdf"],
},
provider: "anthropic",
maxTokens: 5000,
});

Example 3: Data Visualization Analysis​

Analyze charts and underlying data together:

const dataAnalysis = await neurolink.generate({
input: {
text: "Analyze these sales charts and verify against the raw data",
images: [
"https://example.com/q1-chart.png",
"https://example.com/q2-chart.png",
],
csvFiles: ["./sales-data.csv"],
},
provider: "vertex",
enableAnalytics: true,
enableEvaluation: true,
});

Summary​

NeuroLink's multimodal capabilities provide:

✅ Universal input support - Images, PDFs, CSV files ✅ Provider flexibility - Extensive provider compatibility matrix ✅ Automatic format detection - Smart file type recognition ✅ Accessibility features - Alt text support for images ✅ Production-ready - Battle-tested at enterprise scale ✅ Developer-friendly - Works seamlessly across CLI and SDK

Next Steps:

  1. Review the provider support matrix to select the right provider
  2. Try the quick start examples with your use case
  3. Explore advanced recipes for complex scenarios
  4. Check troubleshooting if you encounter issues