Choosing a Model

With Local AI, you choose the language model that AI Assist uses. There is no single best model: new models appear regularly, and the right choice depends on your computer and the AI Assist features you use most. This page explains what to look for so you can choose a model that runs well on your hardware and produces useful results. To learn how to connect MAXQDA to the model you choose, see setting up Local AI.

Local models compared with MAXQDA's cloud

The main benefit of a local model is control: your data stays on your computer or within your organization's network. The trade-off is that results are usually less accurate and responses are slower than with MAXQDA's cloud, for two main reasons:

  • Size: The models behind MAXQDA's cloud are far larger than the models a single computer, or most organizations, can run.
  • Tuning: AI Assist is tuned for the specific cloud models it uses. Other models behave differently, for example in how closely they follow instructions or how well they handle a particular language.

You can switch between Local AI and MAXQDA's cloud at any time, for example to use the cloud for more demanding tasks.

Which features need a stronger model

AI Assist features vary in how much they demand from a model:

  • Lighter: Explain This and Translate This work well with almost any model.
  • Medium: AI Chat and AI Summaries need a model that follows instructions reliably.
  • Heavier: AI Coding needs the most capable model because it has to apply your code definitions consistently.

Some features also send several requests at once. For example, when you code several documents with AI Coding, MAXQDA's cloud handles the requests in parallel, while a local server usually processes them one after another. As a result, coding ten documents can take about ten times as long as coding one.

Fit to your computer

Memory

Memory is the most important limitation. Choose a model that fits comfortably in your computer's memory (RAM, or the memory of your graphics card), with enough room left for your data, the operating system, and other programs. As a rule of thumb, keep about 20% of your memory free. A model that is too large may still start, but then runs very slowly or stops with an error.

Model size and quantization

Model sizes are given in parameters, for example 7B (7 billion) or 30B. Larger models usually produce better results but require more memory. Most model libraries also offer each model in several compressed versions, called quantizations and labeled Q2 to Q8. Lower numbers require less memory but reduce quality.

  • Choose Q4 or higher. Q4 and Q5 offer a good balance for most computers.
  • Down to Q4, a larger model at a lower quantization usually performs better than a smaller model at a higher one. For example, a 30B model at Q4 typically gives better results than a 15B model at Q8.
  • Below Q4, quality drops quickly. If only a Q2 or Q3 version of a model fits, a smaller model at Q5 or Q6 is usually the better choice.

Some models, called mixture-of-experts models, list two parameter counts: the total number of parameters determines how much memory the model needs, while the number of active parameters affects how fast it runs.

Model age and quality

Open models improve quickly, and a recent small model can outperform an older, larger one. Models of the same size can also differ considerably, for example in how well they handle different languages. Before you decide, compare models using independent benchmarks such as Artificial Analysis.

Your hardware

  • Graphics card: A graphics card with plenty of its own memory, for example 12 GB or more, provides the fastest responses.
  • Apple silicon Macs: These Macs share a large pool of memory between the processor and graphics chip, so they can often run larger models than a comparable laptop.
  • Processor only: Models also run without a suitable graphics card, but much more slowly. Expect to use smaller models, around 7B to 9B, and wait longer for results.

Reasoning models

Some models "think" before they answer. This can help with complex problems, but it also adds considerable waiting time for tasks such as summarizing, translating, or chatting. Smaller models can also get stuck thinking without producing an answer. For AI Assist, choose a model without reasoning, or set reasoning to low if the model offers that option.

If a request doesn't work

A failed request isn't always a sign that the model is too weak. Language models include an element of randomness, so the same request may succeed on a second try. If a request fails:

  • Try again: Send the same request once more.
  • Adjust the request: Simplify your instructions or select less data. If your data is too large for the model, you can also choose a model that can handle more text at once, meaning one with a larger context window.
  • Try a different model: A different model, not necessarily a larger one, may handle the request better.

Checklist for choosing a model

  1. Check how much memory your computer has and leave about 20% free.
  2. Choose the largest recent model that fits at Q4 or higher. If you use Jan, it suggests models that suit your hardware.
  3. Prefer a model without reasoning, or set reasoning to low.
  4. Try it with a lighter feature first, such as Explain This, and then with the features you use most.
  5. If results are too slow or not good enough, try a different model or quantization.

Was this page helpful?