Cyllama API Reference¶
Complete API reference for cyllama, a high-performance Python library for LLM inference built on llama.cpp.
Table of Contents¶
High-Level Generation API¶
The high-level API provides simple, Pythonic functions and classes for text generation.
complete()¶
One-shot text generation function.
def complete(
prompt: str,
model_path: str,
config: Optional[GenerationConfig] = None,
stream: bool = False,
**kwargs
) -> Response | Iterator[str]
Parameters:
-
prompt(str): Input text prompt -
model_path(str): Path to GGUF model file -
config(GenerationConfig, optional): Generation configuration object -
stream(bool): If True, return iterator of text chunks -
**kwargs: Override config parameters (temperature, max_tokens, etc.)
Returns:
-
Response: Response object with text and stats (if stream=False) -
Iterator[str]: Iterator of text chunks (if stream=True)
Example:
from cyllama import complete
response = complete(
"What is Python?",
model_path="models/llama.gguf",
temperature=0.7,
max_tokens=200
)
# Streaming
for chunk in complete("Tell me a story", model_path="models/llama.gguf", stream=True):
print(chunk, end="", flush=True)
chat()¶
Chat-style generation with message history. Automatically applies the model's built-in chat template.
def chat(
messages: List[Dict[str, str]],
model_path: str,
config: Optional[GenerationConfig] = None,
stream: bool = False,
template: Optional[str] = None,
**kwargs
) -> str | Iterator[str]
Parameters:
-
messages(List[Dict]): List of message dicts with 'role' and 'content' keys -
model_path(str): Path to GGUF model file -
config(GenerationConfig, optional): Generation configuration -
stream(bool): Enable streaming output -
template(str, optional): Chat template name to use. If None, uses model's default. -
**kwargs: Override config parameters
Returns:
-
Response: Response object with text and stats (if stream=False) -
Iterator[str]: Iterator of text chunks (if stream=True)
Example:
from cyllama import chat
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is machine learning?"}
]
response = chat(messages, model_path="models/llama.gguf")
# With explicit template
response = chat(messages, model_path="models/llama.gguf", template="chatml")
apply_chat_template()¶
Apply a chat template to format messages into a prompt string.
def apply_chat_template(
messages: List[Dict[str, str]],
model_path: str,
template: Optional[str] = None,
add_generation_prompt: bool = True,
verbose: bool = False,
) -> str
Parameters:
-
messages(List[Dict]): List of message dicts with 'role' and 'content' keys -
model_path(str): Path to GGUF model file -
template(str, optional): Template name or string. If None, uses model's default. -
add_generation_prompt(bool): Add assistant prompt prefix (default: True) -
verbose(bool): Enable detailed logging
Returns:
str: Formatted prompt string
Supported Templates:
-
llama2, llama3, llama4
-
chatml (Qwen, Yi, etc.)
-
mistral-v1, mistral-v3, mistral-v7
-
phi3, phi4
-
deepseek, deepseek2, deepseek3
-
gemma, falcon3, command-r, vicuna, zephyr, and more
Example:
from cyllama.api import apply_chat_template
messages = [
{"role": "system", "content": "You are helpful."},
{"role": "user", "content": "Hello!"}
]
prompt = apply_chat_template(messages, "models/llama.gguf")
print(prompt)
# <|begin_of_text|><|start_header_id|>system<|end_header_id|>
# You are helpful.<|eot_id|><|start_header_id|>user<|end_header_id|>
# Hello!<|eot_id|><|start_header_id|>assistant<|end_header_id|>
get_chat_template()¶
Get the chat template string from a model.
Parameters:
-
model_path(str): Path to GGUF model file -
template_name(str, optional): Specific template name to retrieve
Returns:
str: Template string (Jinja-style), or empty string if not found
Example:
from cyllama.api import get_chat_template
template = get_chat_template("models/llama.gguf")
print(template) # Shows the Jinja-style template
Response Class¶
Structured response object returned by generation functions.
@dataclass
class Response:
text: str # Generated text content
stats: Optional[GenerationStats] # Generation statistics
finish_reason: str = "stop" # Why generation stopped
model: str = "" # Model path used
Attributes:
-
text(str): The generated text content -
stats(GenerationStats, optional): Statistics including timing and token counts -
finish_reason(str): Reason for completion ("stop", "length", etc.) -
model(str): Path to the model used
String Compatibility:
Response implements the string protocol for backward compatibility:
-
str(response)returnsresponse.text -
response == "string"compares with text -
len(response)returns text length -
for char in response:iterates over text characters -
"substring" in responsechecks text containment -
response + " more"concatenates text
Methods:
to_dict()¶
Convert response to dictionary.
to_json()¶
Convert response to JSON string.
Example:
from cyllama import complete
response = complete("What is Python?", model_path="model.gguf")
# Use as string (backward compatible)
print(response) # Prints text
if "programming" in response:
print("Mentioned programming!")
# Access structured data
print(f"Finish reason: {response.finish_reason}")
if response.stats:
print(f"Tokens/sec: {response.stats.tokens_per_second:.1f}")
# Serialize
data = response.to_dict()
json_str = response.to_json(indent=2)
GenerationStats Class¶
Statistics from a generation run.
@dataclass
class GenerationStats:
prompt_tokens: int # Number of tokens in prompt
generated_tokens: int # Number of tokens generated
total_time: float # Total generation time (seconds)
tokens_per_second: float # Generation speed
prompt_time: float # Time for prompt processing
generation_time: float # Time for token generation
LLM Class¶
Reusable generator with model caching for improved performance.
class LLM:
def __init__(
self,
model_path: str,
config: Optional[GenerationConfig] = None,
verbose: bool = False
)
Parameters:
-
model_path(str): Path to GGUF model file -
config(GenerationConfig, optional): Default generation configuration -
verbose(bool): Print detailed information during generation
Methods:
__call__()¶
Generate text from a prompt.
def __call__(
self,
prompt: str,
config: Optional[GenerationConfig] = None,
stream: bool = False,
on_token: Optional[Callable[[str], None]] = None
) -> Response | Iterator[str]
Parameters:
-
prompt(str): Input text -
config(GenerationConfig, optional): Override instance config -
stream(bool): Enable streaming -
on_token(Callable, optional): Callback for each token
Returns:
-
Response: Response object with text and stats (if stream=False) -
Iterator[str]: Iterator of text chunks (if stream=True)
chat()¶
Generate a response from chat messages using the model's chat template.
def chat(
self,
messages: List[Dict[str, str]],
config: Optional[GenerationConfig] = None,
stream: bool = False,
template: Optional[str] = None
) -> str | Iterator[str]
Parameters:
-
messages(List[Dict]): List of message dicts with 'role' and 'content' keys -
config(GenerationConfig, optional): Override instance config -
stream(bool): Enable streaming -
template(str, optional): Chat template name to use
get_chat_template()¶
Get the chat template string from the loaded model.
Example:
from cyllama import LLM, GenerationConfig
gen = LLM("models/llama.gguf")
# Simple generation
response = gen("What is Python?")
# With custom config
config = GenerationConfig(temperature=0.9, max_tokens=100)
response = gen("Tell me a joke", config=config)
# With statistics
response, stats = gen.generate_with_stats("Question?")
print(f"Generated {stats.generated_tokens} tokens in {stats.total_time:.2f}s")
print(f"Speed: {stats.tokens_per_second:.2f} tokens/sec")
# Chat with template
messages = [{"role": "user", "content": "Hello!"}]
response = gen.chat(messages)
# Get template
template = gen.get_chat_template()
MCP client methods¶
Since 0.2.11 LLM can attach to Model Context Protocol servers and drive a tool-calling loop against their tools:
def add_mcp_server(
self,
name: str,
*,
command: Optional[str] = None,
args: Optional[list[str]] = None,
env: Optional[dict[str, str]] = None,
cwd: Optional[str] = None,
url: Optional[str] = None,
headers: Optional[dict[str, str]] = None,
transport: Optional["McpTransportType"] = None,
request_timeout: Optional[float] = None,
shutdown_timeout: Optional[float] = None,
) -> None
def remove_mcp_server(self, name: str) -> None
def list_mcp_tools(self) -> list["McpTool"]
def list_mcp_resources(self) -> list["McpResource"]
def call_mcp_tool(self, name: str, arguments: dict) -> Any
def read_mcp_resource(self, uri: str) -> str
def chat_with_tools(
self,
messages: list[dict],
*,
tools: Optional[list["Tool"]] = None,
use_mcp: bool = True,
max_iterations: int = 8,
verbose: bool = False,
system_prompt: Optional[str] = None,
generation_config: Optional[GenerationConfig] = None,
) -> str
See MCP Client for stdio/HTTP quick-start, per-method semantics, and examples of mixing local Tools with MCP tools.
GenerationConfig Dataclass¶
Configuration for text generation.
@dataclass
class GenerationConfig:
max_tokens: int = 512
temperature: float = 0.8
top_k: int = 40
top_p: float = 0.95
min_p: float = 0.05
repeat_penalty: float = 1.0
penalty_last_n: int = 64
frequency_penalty: float = 0.0
presence_penalty: float = 0.0
dry_multiplier: float = 0.0
dry_base: float = 1.75
dry_allowed_length: int = 2
dry_penalty_last_n: int = -1
dry_sequence_breakers: List[str] = ["\n", ":", '"', "*"]
top_n_sigma: float = -1.0
mirostat: int = 0
mirostat_tau: float = 5.0
mirostat_eta: float = 0.1
n_gpu_layers: int = -1
main_gpu: int = 0
split_mode: int = 1
tensor_split: Optional[List[float]] = None
n_ctx: Optional[int] = None
n_batch: int = 2048
n_threads: int = -1
n_threads_batch: int = -1
seed: int = LLAMA_DEFAULT_SEED # 0xFFFFFFFF
stop_sequences: List[str] = field(default_factory=list)
add_bos: bool = True
parse_special: bool = True
Defaults live in cyllama/defaults.py.
Attributes:
-
max_tokens: Maximum tokens to generate (default: 512) -
temperature: Sampling temperature, 0.0 = greedy (default: 0.8) -
top_k: Top-k sampling parameter (default: 40) -
top_p: Top-p (nucleus) sampling (default: 0.95) -
min_p: Minimum probability threshold (default: 0.05) -
repeat_penalty: Penalty for repeating tokens (default: 1.0, disabled) -
penalty_last_n: Number of recent tokens considered for penalties; 0 = disabled, -1 = full context (default: 64) -
frequency_penalty: Penalize tokens by frequency in the recent window, 0.0 = disabled (default: 0.0) -
presence_penalty: Penalize tokens already present in the recent window, 0.0 = disabled (default: 0.0) -
dry_multiplier: DRY repetition penalty scale, 0.0 = disabled (default: 0.0). DRY penalizes tokens that would extend a phrase already in the context;repeat_penaltyworks per token. -
dry_base: Exponential base for DRY penalty growth (default: 1.75) -
dry_allowed_length: Repetitions up to this length go unpenalized (default: 2) -
dry_penalty_last_n: Tokens scanned for repetitions; 0 = disabled, -1 = full context (default: -1) -
dry_sequence_breakers: Strings that reset DRY's repetition tracking (default:["\n", ":", '"', "*"]) -
top_n_sigma: Keep tokens within n standard deviations of the top logit; -1.0 = disabled (default: -1.0). When enabled it replaces top_k / top_p / min_p. -
mirostat: Mirostat sampling mode -- 0 = off, 1 = v1, 2 = v2. When enabled, replaces top_k / top_p / min_p / temperature with the mirostat sampler (default: 0) -
mirostat_tau: Mirostat target entropy (default: 5.0) -
mirostat_eta: Mirostat learning rate (default: 0.1) -
n_gpu_layers: GPU layers to offload (default: -1 = all) -
main_gpu: Primary GPU device index (default: 0) -
split_mode: Multi-GPU split -- 0 = none (main_gpuonly), 1 = layers and KV cache, 2 = rows / tensor parallelism (default: 1) -
tensor_split: Proportion of work per GPU, normalized by llama.cpp;[1, 2]gives GPU 1 two thirds (default: None = auto) -
n_ctx: Context window size, None = prompt length +max_tokens(default: None) -
n_batch: Maximum tokens perllama_decodecall during prompt processing (default: 2048) -
n_threads: CPU threads for generation; -1 = physical cores (default: -1). Logical cores (SMT) slow generation, which is memory-bound. -
n_threads_batch: CPU threads for prompt processing; -1 = physical cores (default: -1). Prompt processing is compute-bound, so logical cores can help when nothing else is decoding. -
seed: Random seed; the defaultLLAMA_DEFAULT_SEED(0xFFFFFFFF) picks a random seed per call and disables the result cache -
stop_sequences: Strings that stop generation (default: []) -
add_bos: Add beginning-of-sequence token (default: True) -
parse_special: Parse special tokens in prompt (default: True)
GenerationStats Dataclass¶
Statistics from a generation run.
@dataclass
class GenerationStats:
prompt_tokens: int
generated_tokens: int
total_time: float
tokens_per_second: float
prompt_time: float = 0.0
generation_time: float = 0.0
Async API¶
The async API provides non-blocking generation for use in async applications (FastAPI, aiohttp, etc.).
AsyncLLM Class¶
Async wrapper around the LLM class for non-blocking text generation.
class AsyncLLM:
def __init__(
self,
model_path: str,
config: Optional[GenerationConfig] = None,
verbose: bool = False,
**kwargs
)
Parameters:
-
model_path(str): Path to GGUF model file -
config(GenerationConfig, optional): Generation configuration -
verbose(bool): Print detailed information during generation -
**kwargs: Generation parameters (temperature, max_tokens, etc.)
Methods:
__call__() / generate()¶
Generate text asynchronously.
stream()¶
Stream generated text chunks asynchronously.
async def stream(
self,
prompt: str,
config: Optional[GenerationConfig] = None,
**kwargs
) -> AsyncIterator[str]
generate_with_stats()¶
Generate text and return statistics.
async def generate_with_stats(
self,
prompt: str,
config: Optional[GenerationConfig] = None
) -> Tuple[str, GenerationStats]
Example:
import asyncio
from cyllama import AsyncLLM
async def main():
# Context manager ensures cleanup
async with AsyncLLM("model.gguf", temperature=0.7) as llm:
# Simple generation
response = await llm("What is Python?")
print(response)
# Streaming
async for chunk in llm.stream("Tell me a story"):
print(chunk, end="", flush=True)
# With stats
text, stats = await llm.generate_with_stats("Question?")
print(f"Generated {stats.generated_tokens} tokens")
asyncio.run(main())
complete_async()¶
Async convenience function for one-off text completion.
async def complete_async(
prompt: str,
model_path: str,
config: Optional[GenerationConfig] = None,
verbose: bool = False,
**kwargs
) -> str
Example:
chat_async()¶
Async convenience function for chat-style generation.
async def chat_async(
messages: List[Dict[str, str]],
model_path: str,
config: Optional[GenerationConfig] = None,
verbose: bool = False,
**kwargs
) -> str
Example:
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is Python?"}
]
response = await chat_async(messages, model_path="model.gguf")
stream_complete_async()¶
Async streaming completion for one-off use.
async def stream_complete_async(
prompt: str,
model_path: str,
config: Optional[GenerationConfig] = None,
verbose: bool = False,
**kwargs
) -> AsyncIterator[str]
Example:
async for chunk in stream_complete_async("Tell me a story", "model.gguf"):
print(chunk, end="", flush=True)
Framework Integrations¶
OpenAI-Compatible API¶
Drop-in replacement for OpenAI Python client.
OpenAICompatibleClient Class¶
from cyllama.integrations.openai_compat import OpenAICompatibleClient
class OpenAICompatibleClient:
def __init__(
self,
model_path: str,
temperature: float = 0.7,
max_tokens: int = 512,
n_gpu_layers: int = -1
)
Attributes:
chat: Chat completions interface
Example:
from cyllama.integrations.openai_compat import OpenAICompatibleClient
client = OpenAICompatibleClient(model_path="models/llama.gguf")
response = client.chat.completions.create(
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is Python?"}
],
temperature=0.7,
max_tokens=200
)
print(response.choices[0].message.content)
# Streaming
for chunk in client.chat.completions.create(
messages=[{"role": "user", "content": "Count to 5"}],
stream=True
):
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
LangChain Integration¶
Full LangChain LLM interface implementation.
CyllamaLLM Class¶
from cyllama.integrations import CyllamaLLM
class CyllamaLLM(LLM):
model_path: str
temperature: float = 0.7
max_tokens: int = 512
top_k: int = 40
top_p: float = 0.95
repeat_penalty: float = 1.0
n_gpu_layers: int = -1
Example:
from cyllama.integrations import CyllamaLLM
from langchain.chains import LLMChain
from langchain.prompts import PromptTemplate
llm = CyllamaLLM(model_path="models/llama.gguf", temperature=0.7)
prompt = PromptTemplate(
input_variables=["topic"],
template="Explain {topic} in simple terms:"
)
chain = LLMChain(llm=llm, prompt=prompt)
result = chain.run(topic="quantum computing")
# With streaming
from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
llm = CyllamaLLM(
model_path="models/llama.gguf",
streaming=True,
callbacks=[StreamingStdOutCallbackHandler()]
)
Memory Utilities¶
Tools for estimating and optimizing GPU memory usage.
estimate_gpu_layers()¶
Estimate how many layers fit in the given GPU memory.
def estimate_gpu_layers(
model_path: Union[str, Path],
gpu_memory_mb: Union[int, List[int]],
ctx_size: int = 2048,
batch_size: int = 1,
n_parallel: int = 1,
kv_cache_type: str = "f16",
use_mmap: bool = True,
verbose: bool = False,
) -> MemoryEstimate
Parameters:
-
model_path: Path to GGUF model file -
gpu_memory_mb: Available GPU memory in MB; a list for multiple GPUs -
ctx_size: Context size -
batch_size: Batch size -
n_parallel: Number of parallel sequences -
kv_cache_type: KV cache precision,"f16"or"f32"
Returns:
MemoryEstimate. Invalid input logs an error and returns an estimate with every field 0.
Example:
from cyllama import estimate_gpu_layers
estimate = estimate_gpu_layers("models/llama.gguf", gpu_memory_mb=8000, ctx_size=4096)
print(f"GPU layers: {estimate.layers}")
print(f"KV cache: {estimate.vram_kv / 1024**2:.0f} MB")
estimate_memory_usage()¶
Estimate memory needs without a GPU budget.
def estimate_memory_usage(
model_path: Union[str, Path],
ctx_size: int = 2048,
batch_size: int = 1,
verbose: bool = False,
) -> Dict[str, Any]
Returns: a dict. Invalid input returns {"error": "..."}.
{
"model_size_mb": {"f32": 4074, "f16": 2037, "q4_0": 509, "q8_0": 1018},
"kv_cache_mb": {"f16": 256, "f32": 512},
"graph_mb": 1471,
"parameters": {"n_embd": 2048, "n_layer": 16, "n_ff": 8192, "n_vocab": 128256, "total_params": 1067974656},
}
MemoryEstimate Dataclass¶
Result of estimate_gpu_layers().
@dataclass
class MemoryEstimate:
layers: int # layers to offload to GPU
graph_size: int # computation graph (bytes)
vram: int # GPU memory budget used (bytes)
vram_kv: int # KV cache for the offloaded layers (bytes)
total_size: int # weights of all layers (bytes)
tensor_split: Optional[List[int]] = None # layers per GPU, multi-GPU only
Core llama.cpp API¶
Low-level Cython wrappers for direct llama.cpp access.
Core Classes¶
LlamaModel¶
Represents a loaded GGUF model.
from cyllama.llama.llama_cpp import (
LLAMA_LOAD_MODE_MMAP,
LlamaModel,
LlamaModelParams,
)
params = LlamaModelParams()
params.n_gpu_layers = -1
# load_mode: LLAMA_LOAD_MODE_AUTO (default) / _NONE / _MMAP / _MLOCK / _DIRECT_IO
params.load_mode = LLAMA_LOAD_MODE_MMAP
# lazy_mode: LLAMA_LAZY_MODE_OFF / _AUTO (default) / _ON -- on-demand row reads
# for tensors the arch marks; requires mmap
model = LlamaModel("models/llama.gguf", params)
# Properties
print(model.n_params) # Total parameters
print(model.n_layer) # Number of layers
print(model.n_embd) # Embedding dimension
print(model.desc) # Model description
print(model.size) # On-disk size in bytes
# Methods
vocab = model.get_vocab() # Get vocabulary
print(vocab.n_vocab) # Vocabulary size
model.close() # Free resources (or use `with LlamaModel(...) as model:`)
LlamaModel.from_fileobj(fileobj, offset=None, params=None) loads from an open
binary file or an int fd. The GGUF may be embedded in a larger file at offset
(default: the current position). With mmap, the GGUF data section must sit at a
32-byte aligned file offset; use LLAMA_LOAD_MODE_NONE otherwise. The caller's
file position is unchanged.
LlamaContext¶
Inference context for model.
from cyllama.llama.llama_cpp import LlamaContext, LlamaContextParams
ctx_params = LlamaContextParams()
ctx_params.n_ctx = 2048
ctx_params.n_batch = 512
# The C default is 4 threads. The high-level API uses physical cores:
from cyllama.utils.platform import resolve_n_threads
ctx_params.n_threads = ctx_params.n_threads_batch = resolve_n_threads()
ctx = LlamaContext(model, ctx_params)
# Decode batch
from cyllama.llama.llama_cpp import llama_batch_get_one
batch = llama_batch_get_one(tokens)
ctx.decode(batch)
# Memory (KV cache) management
ctx.kv_cache_clear()
ctx.memory_seq_rm(seq_id, p0, p1)
ctx.memory_seq_cp(seq_id_src, seq_id_dst, p0, p1)
ctx.memory_seq_keep(seq_id)
ctx.memory_seq_add(seq_id, p0, p1, delta)
print(ctx.memory_seq_pos_min(seq_id), ctx.memory_seq_pos_max(seq_id))
# Performance
ctx.print_perf_data()
LlamaSampler¶
Sampling strategies for token generation.
from cyllama.llama.llama_cpp import LlamaSampler, LlamaSamplerChainParams
sampler_params = LlamaSamplerChainParams()
sampler = LlamaSampler(sampler_params)
# Add sampling methods
sampler.add_top_k(40)
sampler.add_top_p(0.95, 1)
sampler.add_temp(0.7)
sampler.add_dist(seed)
# Sample token
token_id = sampler.sample(ctx, idx)
# Reset state
sampler.reset()
Backend sampling [EXPERIMENTAL upstream]. A chain attached to a context runs
inside the decode graph on the output device. sample() then returns the backend
token, or finishes on the CPU from the first link that cannot be offloaded.
Offloadable links: greedy, dist, top-k, top-p, min-p, temp, temp-ext, penalties,
logit-bias.
chain = LlamaSampler()
chain.add_top_k(40)
chain.add_temp(0.7)
chain.add_dist(seed)
ctx = LlamaContext(model, ctx_params, samplers={0: chain}) # or ctx.set_sampler(0, chain)
ctx.decode(batch)
token_id = chain.sample(ctx, -1)
ctx.sampled_token_ith(-1) # backend token, or None
ctx.sampled_probs_ith(-1) # aligned with sampled_candidates_ith(-1), or None
ctx.set_sampler(0, None) # detach
Attaching binds the chain to that context for life: llama.cpp never resets the
chain's backend state. A bound chain cannot be attached again, sampled with
another context, or modified, and cannot be closed while attached. Each case
raises instead of aborting or sampling wrong. chain.clone() returns an unbound
copy with the same configuration.
LlamaVocab¶
Vocabulary and tokenization.
vocab = model.get_vocab()
# Tokenization
tokens = vocab.tokenize("Hello world", add_special=True, parse_special=True)
# Detokenization
text = vocab.detokenize(tokens)
piece = vocab.token_to_piece(token_id, special=True)
raw = vocab.token_to_bytes(token_id, special=True) # may be a partial UTF-8 character
# Streaming detokenization: holds a split character until complete
from cyllama.llama.token_decoder import TokenDecoder
decoder = TokenDecoder(vocab)
text = "".join(decoder.decode(t) for t in token_ids) + decoder.flush()
# Special tokens
print(vocab.bos) # Begin-of-sequence token
print(vocab.eos) # End-of-sequence token
print(vocab.eot) # End-of-turn token
print(vocab.n_vocab) # Vocabulary size
# Check token types
is_eog = vocab.is_eog(token_id)
is_control = vocab.is_control(token_id)
LlamaBatch¶
Efficient batch processing.
from cyllama.llama.llama_cpp import LlamaBatch
# Create batch
batch = LlamaBatch(n_tokens=512, embd=0, n_seq_max=1)
# Add token
batch.add(token_id, pos, seq_ids=[0], logits=True)
# Clear batch
batch.clear()
# Convenience function
from cyllama.llama.llama_cpp import llama_batch_get_one
batch = llama_batch_get_one(tokens, pos_offset=0)
Backend Management¶
from cyllama.llama.llama_cpp import (
ggml_backend_load_all,
ggml_backend_reg_names,
ggml_backend_reg_count,
ggml_backend_dev_info,
ggml_backend_unload,
)
# Load all available backends (Metal, CUDA, etc.)
ggml_backend_load_all()
# Which backends registered at runtime
print(ggml_backend_reg_count(), ggml_backend_reg_names()) # e.g. ['CPU', 'Vulkan']
# Per-device information (dicts with name, description, type)
for dev in ggml_backend_dev_info():
print(dev)
# Unload a registered backend by name
ggml_backend_unload("Vulkan")
To ask what was compiled in (as opposed to what loaded), use the build config:
from cyllama._internal import build_config
print(build_config.backend_enabled("cuda"))
print(build_config.backend()) # full per-backend dict
print(build_config.versions()) # pinned llama.cpp / whisper.cpp / sd.cpp revisions
Advanced Features¶
GGUF File Manipulation¶
Inspect and modify GGUF model files.
GGUFContext Class¶
from cyllama.llama.llama_cpp import GGUFContext
# Read existing file
ctx = GGUFContext.from_file("model.gguf")
# From an open binary file or fd, optionally embedded at an offset.
# data_offset is then relative to the start of the containing file.
with open("bundle.bin", "rb") as f:
ctx = GGUFContext.from_fileobj(f, offset=4096)
# Get metadata
metadata = ctx.get_all_metadata()
print(metadata['general.architecture'])
print(metadata['general.name'])
value = ctx.get_value("general.architecture")
# Create new file
ctx = GGUFContext.empty()
ctx.set_val_str("custom.key", "value")
ctx.set_val_u32("custom.number", 42)
ctx.write_to_file("custom.gguf", only_meta=True)
# Modify existing
ctx = GGUFContext.from_file("model.gguf")
ctx.set_val_str("custom.metadata", "updated")
ctx.write_to_file("modified.gguf")
JSON Schema to Grammar¶
Convert JSON schemas to llama.cpp grammar format for structured output. This is implemented in pure Python (vendored from llama.cpp) with no C++ dependency.
from cyllama.llama.llama_cpp import json_schema_to_grammar
schema = {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"},
"email": {"type": "string"}
},
"required": ["name", "age"]
}
grammar = json_schema_to_grammar(schema)
# Use with generation
from cyllama.llama.llama_cpp import LlamaSampler
sampler = LlamaSampler()
sampler.add_grammar(grammar)
Model Download¶
Download models from HuggingFace with Ollama-style tags.
from cyllama.llama.llama_cpp import download_model, list_cached_models
# Download from HuggingFace
download_model(
hf_repo="bartowski/Llama-3.2-1B-Instruct-GGUF:q4",
cache_dir="~/.cache/cyllama/models"
)
# List cached models
models = list_cached_models()
for model in models:
print(f"{model['user']}/{model['model']}:{model['tag']}")
print(f" Path: {model['path']}")
print(f" Size: {model['size'] / 1024 / 1024:.2f} MB")
# Direct URL download
download_model(
url="https://example.com/model.gguf",
output_path="models/custom.gguf"
)
N-gram Cache¶
Pattern-based token prediction for 2-10x speedup on repetitive text.
from cyllama.llama.llama_cpp import NgramCache
# Create cache
cache = NgramCache()
# Learn patterns from token sequences
tokens = [1, 2, 3, 4, 5, 6, 7, 8]
cache.update(tokens, ngram_min=2, ngram_max=4)
# Predict likely continuations
input_tokens = [1, 2, 3]
draft_tokens = cache.draft(input_tokens, n_draft=16)
# Save/load cache
cache.save("patterns.bin")
loaded_cache = NgramCache.from_file("patterns.bin")
# Clear cache
cache.clear()
LoRA Adapters¶
Apply LoRA adapters to a context. Adapters are a binding-layer feature: the
high-level LLM does not expose them, because it recreates its context when a
prompt needs a larger one and adapters would silently stop applying.
from cyllama.llama.llama_cpp import LlamaModel, LlamaContext, LlamaContextParams
model = LlamaModel("models/llama.gguf")
# An adapter is owned by the model that loads it and stays valid for that
# model's lifetime. The adapter holds a reference to its model, so the model
# cannot be collected while any adapter borrowed from it is still alive.
adapter = model.lora_adapter_init("models/adapter.gguf")
ctx_params = LlamaContextParams()
ctx_params.n_ctx = 2048
ctx = LlamaContext(model, ctx_params)
# Apply at scale 1.0. Model weights are not modified.
ctx.set_adapters_lora([(adapter, 1.0)])
# Several adapters at once, as pairs or as a mapping.
second = model.lora_adapter_init("models/other.gguf")
ctx.set_adapters_lora({adapter: 0.8, second: 0.5})
# Empty clears every adapter.
ctx.set_adapters_lora()
set_adapters_lora() replaces the whole set on each call rather than adding to
it, matching the llama.cpp call it wraps. A scale of 0.0 drops an adapter.
Passing the same adapter twice raises ValueError, because llama.cpp keys its
set by pointer and would otherwise keep only one of the two scales.
Methods:
| Method | Description |
|---|---|
LlamaModel.lora_adapter_init(path) |
Load an adapter against this model |
LlamaModel.lora_adapter_init_from_fileobj(fileobj, offset=None) |
Same, from an open binary file or int fd, optionally at an offset |
LlamaContext.set_adapters_lora(adapters=()) |
Replace the context's adapter set |
LlamaContext.lora_adapters |
Adapters currently set, in the order given |
LlamaAdapterLora.model |
The model that owns this adapter |
LlamaAdapterLora.meta_count() |
Number of GGUF metadata key/value pairs |
LlamaAdapterLora.meta_val_str(key) |
Metadata value by key name |
LlamaAdapterLora.meta_key_by_index(i) |
Metadata key name by index |
LlamaAdapterLora.meta_val_str_by_index(i) |
Metadata value by index |
LlamaAdapterLora.n_alora_invocation_tokens |
Length of the aLoRA invocation sequence, 0 for a plain LoRA |
LlamaAdapterLora.alora_invocation_tokens |
Token sequence that activates an aLoRA |
An activated LoRA (aLoRA) only takes effect once its invocation tokens appear in
the prompt. alora_invocation_tokens is empty for a plain LoRA.
From the CLI, --lora and --lora-scaled do the same thing and are repeatable:
python -m cyllama.llama.cli -m models/llama.gguf \
--lora models/adapter.gguf \
--lora-scaled models/other.gguf 0.5 \
-p "Hello"
Speculative Decoding¶
Use draft model for 2-3x inference speedup.
from cyllama.llama.llama_cpp import (
LlamaModel, LlamaContext, LlamaModelParams, LlamaContextParams,
Speculative, SpeculativeParams
)
# Load target and draft models
model_target = LlamaModel("models/large.gguf", LlamaModelParams())
model_draft = LlamaModel("models/small.gguf", LlamaModelParams())
ctx_params = LlamaContextParams()
ctx_params.n_ctx = 2048
ctx_target = LlamaContext(model_target, ctx_params)
# Configure speculative parameters
params = SpeculativeParams(
n_max=3, # Maximum number of draft tokens
n_min=0, # Minimum number of draft tokens
p_min=0.0 # Minimum acceptance probability
)
# Check target-context compatibility (static method)
if Speculative.is_compat(ctx_target):
print("Target context is compatible for speculative decoding")
ctx_draft = LlamaContext(model_draft, ctx_params)
# Create speculative decoding instance
spec = Speculative(params, ctx_target, ctx_draft)
# Begin a speculative decoding round
prompt_tokens = [1, 2, 3] # tokens the target has processed
spec.begin(prompt_tokens)
# Draft continuations of the target's newest sampled token,
# which is not yet in prompt_tokens
last_token = 4
draft_tokens = spec.draft(params, prompt_tokens, last_token)
# Accept verified tokens (n_accepted from target verification)
spec.accept(n_accepted=len(draft_tokens))
# Print performance statistics
spec.print_stats()
Parameters:
-
n_max: Maximum number of tokens to draft (default: 3) -
n_min: A shorter draft is discarded (default: 0) -
p_min: Drafting stops when the top candidate's probability, a softmax over the draft model's top 10 logits, falls below this (default: 0.0)
Methods:
| Method | Description |
|---|---|
Speculative.is_compat(ctx_target) |
Static: check if target context supports speculative decoding |
begin(prompt_tokens) |
Begin a speculative decoding round |
draft(params, prompt_tokens, last_token_id) |
Draft greedy continuations of last_token_id, the target's newest token, which follows prompt_tokens |
accept(n_accepted) |
Accept the first n_accepted verified draft tokens |
print_stats() |
Print speculative decoding performance statistics |
tests/examples/speculative_example.py shows the full draft/verify loop around
these calls, with --bench to compare against plain greedy decoding.
Server Implementations¶
Three OpenAI-compatible server implementations, all exported from cyllama.llama.server. Note that the package-level ServerConfig is an alias for PythonServerConfig (the dataclass in .python), which the embedded server also takes; the subprocess launcher has its own LauncherServerConfig.
PythonServer¶
Pure Python server implementation.
from cyllama.llama.server import start_python_server
server = start_python_server(
model_path="models/llama.gguf",
host="127.0.0.1",
port=8000,
n_ctx=2048,
n_gpu_layers=-1,
)
# Use with OpenAI client
import openai
client = openai.OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "Hello!"}],
)
server.stop()
EmbeddedServer¶
In-process server built on cpp-httplib. It serves from start() until stop() on a thread pool, and streams chat completions as server-sent events when the request sets "stream": true. It is a compiled extension, so it is only importable when the wheel was built with it -- cyllama.llama.server swallows the ImportError otherwise.
from cyllama.llama.server import EmbeddedServer, ServerConfig
config = ServerConfig(
model_path="models/llama.gguf",
host="127.0.0.1",
port=8080,
n_ctx=2048,
n_threads=4,
)
# Context manager starts the server and stops it on exit
with EmbeddedServer(config) as server:
... # server runs at http://127.0.0.1:8080
# Or the convenience helper, which builds the config and starts it for you
from cyllama.llama.server import start_embedded_server
server = start_embedded_server("models/llama.gguf", port=8080, n_ctx=2048)
server.stop()
LlamaServer¶
Python wrapper around the llama.cpp server binary.
from cyllama.llama.server import LlamaServer, LauncherServerConfig
config = LauncherServerConfig(
model_path="models/llama.gguf",
host="127.0.0.1",
port=8080
)
server = LlamaServer(config, server_binary="bin/llama-server")
server.start()
# Check status
if server.is_running():
print("Server is running")
server.stop()
There is also start_server(model_path, **kwargs) as a one-call shorthand, and LlamaServerClient for talking to a running server:
from cyllama.llama.server import LlamaServerClient
client = LlamaServerClient("http://127.0.0.1:8080")
print(client.health())
print(client.chat_completion([{"role": "user", "content": "Hello!"}]))
Multimodal Support¶
Vision- and audio-language models, via llama.cpp's libmtmd. The high-level classes live in cyllama.llama.mtmd; the underlying MtmdContext, MtmdBitmap, and MtmdInputChunks bindings are re-exported from the same package.
from cyllama.llama.llama_cpp import LlamaModel
from cyllama.llama.mtmd import MultimodalProcessor
model = LlamaModel("models/model.gguf")
# The projector is loaded by the processor itself
processor = MultimodalProcessor("models/mmproj.gguf", model)
print(processor.supports_vision)
print(processor.supports_audio)
# Tokenize text + image into chunks ready for evaluation.
# The media marker is appended automatically if the text lacks one.
chunks = processor.process_image("What's in this image?", "image.jpg")
# Audio works the same way when the model supports it
chunks = processor.process_audio("Transcribe this", "audio.wav")
For a full question-answer loop, VisionLanguageChat wraps the processor and a context and keeps conversation history:
from cyllama.llama.mtmd import VisionLanguageChat
chat = VisionLanguageChat("models/mmproj.gguf", model, ctx)
print(chat.ask_about_image("What's in this image?", "image.jpg"))
print(chat.continue_conversation("What colour is the car?"))
chat.clear_history()
AudioProcessor and ImageAnalyzer provide the equivalent task-specific helpers. All of them raise UnsupportedModalityError (a MultimodalError subclass) when the loaded model lacks the modality.
Whisper Integration¶
Speech-to-text transcription using whisper.cpp. See Whisper.cpp Integration for complete documentation.
Quick Start¶
from cyllama.whisper import WhisperContext, WhisperFullParams
import numpy as np
# Load model
ctx = WhisperContext("models/ggml-base.en.bin")
# Audio must be 16kHz mono float32
samples = load_audio_as_float32("audio.wav") # Your audio loading function
# Transcribe
params = WhisperFullParams()
params.language = "en"
ctx.full(samples, params)
# Get results
for i in range(ctx.full_n_segments()):
t0 = ctx.full_get_segment_t0(i) / 100.0 # centiseconds to seconds
t1 = ctx.full_get_segment_t1(i) / 100.0
text = ctx.full_get_segment_text(i)
print(f"[{t0:.2f}s - {t1:.2f}s] {text}")
Key Classes¶
| Class | Description |
|---|---|
WhisperContext |
Main context for model loading and inference |
WhisperContextParams |
Configuration for context creation |
WhisperFullParams |
Configuration for transcription |
WhisperVadParams |
Voice activity detection parameters |
WhisperContext Methods¶
| Method | Description |
|---|---|
full(samples, params) |
Run transcription on float32 audio samples |
full_n_segments() |
Get number of transcribed segments |
full_get_segment_text(i) |
Get text of segment i |
full_get_segment_t0(i) |
Get start time (centiseconds) |
full_get_segment_t1(i) |
Get end time (centiseconds) |
full_lang_id() |
Get detected language ID |
is_multilingual() |
Check if model supports multiple languages |
Audio Requirements¶
-
Sample rate: 16000 Hz
-
Channels: Mono
-
Format: Float32 normalized to [-1.0, 1.0]
Stable Diffusion Integration¶
Image generation using stable-diffusion.cpp. Supports SD 1.x/2.x, SDXL, SD3, FLUX, video generation (Wan/CogVideoX), and ESRGAN upscaling.
Note: Build with WITH_STABLEDIFFUSION=1 to enable this module.
The module is exposed as cyllama.sd (CLI: python -m cyllama.sd). For broader narrative documentation, see docs/stable_diffusion.md; this section is the API reference.
Quick Start¶
from cyllama.sd import text_to_image
# Simple text-to-image generation
image = text_to_image(
model_path="models/sd_xl_turbo_1.0.q8_0.gguf",
prompt="a photo of a cute cat",
width=512,
height=512,
sample_steps=4,
cfg_scale=1.0
)
# text_to_image returns a single SDImage; text_to_images returns a List[SDImage]
image.save("output.png")
text_to_image()¶
Convenience function that creates a context, generates one image, and tears the context down. Returns a single SDImage. For batches use text_to_images().
def text_to_image(
model_path: str,
prompt: str,
negative_prompt: str = "",
width: int = 512,
height: int = 512,
seed: int = -1,
sample_steps: int = 20,
cfg_scale: float = 7.0,
sample_method: SampleMethod = SampleMethod.COUNT,
scheduler: Scheduler = Scheduler.COUNT,
n_threads: int = -1,
vae_path: Optional[str] = None,
taesd_path: Optional[str] = None,
clip_l_path: Optional[str] = None,
clip_g_path: Optional[str] = None,
t5xxl_path: Optional[str] = None,
control_net_path: Optional[str] = None,
clip_skip: int = -1,
eta: float = float('inf'),
slg_scale: float = 0.0,
vae_tiling: bool = False,
hires_fix: bool = False,
hires_scale: float = 2.0,
diffusion_flash_attn: bool = False
) -> SDImage
SampleMethod.COUNT and Scheduler.COUNT are auto-detect sentinels — the C library picks based on the loaded model. eta=float('inf') resolves to a method-specific default. hires_fix=True enables hires-fix two-pass generation with default latent upscale; for finer control use SDImageGenParams.set_hires_fix(...).
text_to_images()¶
Same as text_to_image() but returns List[SDImage] and accepts batch_count: int = 1. Each image in the batch uses an incremented seed, producing variants of the same prompt.
image_to_image()¶
Img2img convenience function.
def image_to_image(
model_path: str,
init_image: Union[SDImage, str],
prompt: str,
negative_prompt: str = "",
strength: float = 0.75,
seed: int = -1,
sample_steps: int = 20,
cfg_scale: float = 7.0,
sample_method: SampleMethod = SampleMethod.COUNT,
scheduler: Scheduler = Scheduler.COUNT,
n_threads: int = -1,
vae_path: Optional[str] = None,
clip_skip: int = -1
) -> List[SDImage]
init_image accepts either an SDImage or a filesystem path; output dimensions are taken from the init image.
SDContext¶
Persistent generation context — load the model once, generate many times.
from cyllama.sd import SDContext, SDContextParams, SampleMethod, Scheduler
params = SDContextParams()
params.model_path = "models/sd_xl_turbo_1.0.q8_0.gguf"
params.n_threads = 4
with SDContext(params) as ctx:
images = ctx.generate(
prompt="a beautiful landscape",
negative_prompt="blurry, ugly",
width=512, height=512,
sample_steps=4, cfg_scale=1.0,
sample_method=SampleMethod.EULER, # or COUNT for auto-detect
scheduler=Scheduler.DISCRETE,
hires_fix=False,
)
SDContext.generate(...) accepts the same kwargs as text_to_image() plus batch_count, init_image, mask_image, control_image, control_strength, strength, and flow_shift. Returns List[SDImage].
Properties:
-
is_valid(bool): Context loaded successfully. -
supports_image_generation(bool): Model can rungenerate()(false for video-only models). -
supports_video_generation(bool): Model can rungenerate_video().
Methods:
-
generate(**kwargs) -> List[SDImage]: Text/img2img/inpaint/ControlNet generation. -
generate_with_params(params: SDImageGenParams) -> List[SDImage]: Low-level entry point taking a fully populated params object — needed for advanced features (LoRAs, reference images, Photo Maker, PuLID, hires-fix model upscalers, full cache configuration). -
generate_video(**kwargs) -> List[SDImage]: Video frame generation (requires video-capable model). -
default_sample_method(sample_method=None) -> SampleMethod: Model's preferred sampler. -
default_scheduler(sample_method=None) -> Scheduler: Model's preferred scheduler. -
cancel(mode: CancelMode = CancelMode.ALL) -> None: Request cancellation of an in-flightgenerate()/generate_video()running on another thread.CancelMode.ALLstops as soon as possible,CancelMode.NEW_LATENTSfinishes the current sample then skips remaining batch latents,CancelMode.RESETclears a pending request.
SDContextParams¶
Configuration for model loading.
params = SDContextParams()
params.model_path = "model.gguf" # Main model
params.vae_path = "vae.safetensors" # Optional VAE
params.taesd_path = "taesd.safetensors" # Optional TAESD (fast previews)
params.clip_l_path = "clip_l.safetensors" # Optional CLIP-L (SDXL/SD3)
params.clip_g_path = "clip_g.safetensors" # Optional CLIP-G (SDXL/SD3)
params.t5xxl_path = "t5xxl.safetensors" # Optional T5-XXL (SD3/FLUX)
params.control_net_path = "cn.safetensors" # Optional ControlNet
params.n_threads = 4
params.diffusion_flash_attn = False
params.max_vram = "-1" # Per-device budget: "N" GiB cap, "-N" leave N free
params.backend = None # Compute placement, e.g. "diffusion=cuda0,te=cpu"
params.params_backend = "te=cpu" # Weight placement; also accepts "cpu"/"disk"
params.auto_fit = True # Tiered placement by free memory (default)
params.eager_load = False # Load all params up front
params.disable_prefetch = False # Async prefetch of next segment's weights
params.disable_segmented_compute = False # True forces monolithic graphs
params.tokenizer = None # tokenizer.json; required for PiD and Lens
params.pulid_weights_path = None # Optional PuLID weights
params.rpc_servers = None # Optional RPC backends, e.g. "host:port"
params.wtype = SDType.COUNT # COUNT = auto-detect
params.rng_type = RngType.CUDA
Memory is controlled by two independent mechanisms, which together replace the removed offload_params_to_cpu / keep_clip_on_cpu / keep_vae_on_cpu / keep_control_net_on_cpu / free_params_immediately / vae_decode_only flags:
-
Placement --
backendassigns compute per module andparams_backendassigns where the weights live (cpuanddiskare valid targets for the latter). Both take either a bare target for every module ("cuda0","cpu") or comma-separated per-module assignments ("diffusion=cuda0,te=cpu"). Module keys:diffusion(aliasesmodel/unet/dit),te(aliasesclip/text/conditioner/llm/t5),vae,clip-vision,control-net,photomaker,upscaler,detector.params_backend = "te=cpu"is the direct replacement for the oldkeep_clip_on_cpu. -
Budget --
max_vramis a per-device GiB budget shared by resident weights and compute buffers."N"caps each device at N GiB,"-N"leaves N GiB free, and"0"orNoneuses live free VRAM. The per-device form is"cuda0=6,vulkan0=4". A graph that does not fit is split into segments;disable_segmented_compute = Trueprevents that.
auto_fit (default True) places each module on the compute GPU, then RAM, another GPU, or disk, by available memory. A non-empty params_backend disables it.
SDImage¶
Image wrapper with numpy and PIL integration.
from cyllama.sd import SDImage
import numpy as np
arr = np.zeros((512, 512, 3), dtype=np.uint8)
img = SDImage.from_numpy(arr)
print(img.width, img.height, img.channels)
arr = img.to_numpy() # (H, W, C) uint8
pil_img = img.to_pil() # requires Pillow
img.save("output.png")
img = SDImage.load("input.png")
SDImageGenParams¶
Full generation parameters; pass to SDContext.generate_with_params(). The text_to_image() convenience function only exposes a curated subset — drop down to this class for LoRAs, reference images, Photo Maker, PuLID, full cache control, hires-fix model upscalers, etc.
from cyllama.sd import SDImageGenParams, SDImage, HiresUpscaler
params = SDImageGenParams()
params.prompt = "a cute cat"
params.negative_prompt = "ugly, blurry"
params.width = 512
params.height = 512
params.seed = 42
params.batch_count = 1
params.strength = 0.75 # For img2img
params.clip_skip = -1
# VAE tiling
params.vae_tiling_enabled = True
params.vae_tile_size = (512, 512)
params.vae_tile_overlap = 0.5
# Cache acceleration (legacy easycache_* aliases also available)
params.cache_mode = 1 # 0=disabled, 1=easycache, 2=ucache, 3=dbcache, 4=taylorseer, 5=cache_dit
params.cache_threshold = 0.1
params.cache_range = (0.0, 1.0)
# Hires-fix two-pass generation
params.set_hires_fix(
enabled=True,
upscaler=HiresUpscaler.LATENT, # or LANCZOS, NEAREST, MODEL, ...
scale=2.0,
denoising_strength=0.7,
)
# ...individual setters also work:
# params.hires_enabled = True
# params.hires_target_size = (1024, 1024)
# params.hires_model_path = "/path/to/upscaler.gguf" # required for HiresUpscaler.MODEL
# img2img / inpaint / ControlNet
params.set_init_image(SDImage.load("input.png"))
params.set_mask_image(SDImage.load("mask.png"))
params.set_control_image(control_img, strength=0.8)
# LoRAs and reference images
params.set_loras([{"path": "lora.safetensors", "multiplier": 0.8}])
params.set_ref_images([ref_img1, ref_img2])
# Sample params (delegated to nested SDSampleParams)
sample = params.sample_params
sample.sample_steps = 20
sample.cfg_scale = 7.0
sample.sample_method = SampleMethod.COUNT
sample.scheduler = Scheduler.COUNT
See docs/stable_diffusion.md for the full property catalog (Photo Maker, PuLID, ControlNet refs, full cache configuration, all hires-fix fields).
SDSampleParams¶
Sampling configuration. Usually accessed as gen_params.sample_params rather than instantiated directly.
from cyllama.sd import SDSampleParams, SampleMethod, Scheduler
params = SDSampleParams()
params.sample_method = SampleMethod.COUNT
params.scheduler = Scheduler.COUNT
params.sample_steps = 20
params.cfg_scale = 7.0
params.eta = float('inf') # inf = method-specific default
params.slg_scale = 0.0 # Skip layer guidance
params.flow_shift = float('inf') # Flow shift (SD3.x / Wan)
Upscaler¶
ESRGAN-based image upscaling.
from cyllama.sd import Upscaler, SDImage
upscaler = Upscaler(
"models/esrgan-x4.bin",
n_threads=4,
direct=False, # direct convolution
tile_size=0, # 0 = default
)
print(f"Factor: {upscaler.upscale_factor}x")
img = SDImage.load("input.png")
upscaled = upscaler.upscale(img) # use model's native factor
upscaled = upscaler.upscale(img, factor=2) # or override
upscaled.save("upscaled.png")
Upscaler is also usable as a context manager (with Upscaler(...) as up:).
convert_model()¶
Convert models between formats / quantize.
from cyllama.sd import convert_model, SDType
convert_model(
input_path="sd-v1-5.safetensors",
output_path="sd-v1-5-q4_0.gguf",
output_type=SDType.Q4_0,
vae_path="vae-ft-mse.safetensors", # optional
tensor_type_rules=None, # optional per-tensor type rules
convert_name=False, # convert tensor names
)
Raises FileNotFoundError if the input is missing, RuntimeError on conversion failure.
canny_preprocess()¶
Canny edge detection for ControlNet conditioning. Modifies the image in place.
from cyllama.sd import SDImage, canny_preprocess
img = SDImage.load("photo.png")
success = canny_preprocess(
img,
high_threshold=0.8,
low_threshold=0.1,
weak=0.5,
strong=1.0,
inverse=False,
)
Callbacks¶
from cyllama.sd import (
set_log_callback,
set_progress_callback,
set_preview_callback,
PreviewMode,
)
# Logging: callback receives (LogLevel, str)
def log_cb(level, text):
print(f'[{level.name}] {text}', end='')
set_log_callback(log_cb)
# Progress: callback receives (step, total_steps, time_seconds)
def progress_cb(step, steps, time_s):
pct = (step / steps) * 100 if steps > 0 else 0
print(f'Step {step}/{steps} ({pct:.1f}%) - {time_s:.2f}s')
set_progress_callback(progress_cb)
# Preview: callback receives (step, frames: List[SDImage], is_noisy: bool)
def preview_cb(step, frames, is_noisy):
for i, frame in enumerate(frames):
frame.save(f"preview_{step}_{i}.png")
set_preview_callback(
preview_cb,
mode=PreviewMode.TAE,
interval=5,
denoised=True,
noisy=False,
)
# Pass None to clear any of them.
set_log_callback(None)
set_progress_callback(None)
set_preview_callback(None)
Enums¶
SampleMethod
-
EULER,EULER_A,HEUN,DPM2,DPMPP2S_A,DPMPP2M,DPMPP2Mv2 -
IPNDM,IPNDM_V,LCM,DDIM_TRAILING,TCD -
RES_MULTISTEP,RES_2S,ER_SDE -
COUNT(auto-detect sentinel)
Scheduler
-
DISCRETE,KARRAS,EXPONENTIAL,AYS,GITS -
SGM_UNIFORM,SIMPLE,SMOOTHSTEP,KL_OPTIMAL,LCM,BONG_TANGENT -
COUNT(auto-detect sentinel)
Prediction
EPS,V,EDM_V,FLOW,FLUX_FLOW,SEFI_FLOW,MINIT2I_FLOW,COUNT
SDType: Data types for model weights / quantization
-
F32,F16,BF16 -
Q4_0,Q4_1,Q5_0,Q5_1,Q8_0,Q8_1 -
Q2_K,Q3_K,Q4_K,Q5_K,Q6_K,Q8_K -
COUNT(auto-detect sentinel)
RngType: STD_DEFAULT, CUDA, CPU
LogLevel: DEBUG, INFO, WARN, ERROR
PreviewMode: NONE, PROJ, TAE, VAE
LoraApplyMode: AUTO, IMMEDIATELY, AT_RUNTIME
HiresUpscaler: hires-fix upscaler modes
-
NONE -
LATENT,LATENT_NEAREST,LATENT_NEAREST_EXACT,LATENT_ANTIALIASED,LATENT_BICUBIC,LATENT_BICUBIC_ANTIALIASED -
LANCZOS,NEAREST -
MODEL(external upscaler model — sethires_model_path)
Utility Functions¶
from cyllama.sd import (
get_num_cores,
get_system_info,
type_name,
sample_method_name,
scheduler_name,
ggml_backend_load_all,
)
ggml_backend_load_all() # call before get_system_info() so GPU backends register
print(f"CPU cores: {get_num_cores()}")
print(get_system_info())
print(type_name(SDType.Q4_0)) # "q4_0"
print(sample_method_name(SampleMethod.EULER)) # "euler"
print(scheduler_name(Scheduler.KARRAS)) # "karras"
CLI Tool¶
# txt2img (alias: generate)
python -m cyllama.sd txt2img \
--model models/sd_xl_turbo_1.0.q8_0.gguf \
--prompt "a beautiful sunset" \
--output sunset.png \
--steps 4 --cfg 1.0
# img2img / inpaint / ControlNet / video
python -m cyllama.sd img2img --model M --init INPUT --prompt "..." --output OUT
python -m cyllama.sd inpaint --model M --init INPUT --mask MASK --prompt "..." --output OUT
python -m cyllama.sd controlnet --model M --control-net CN --control-image C --prompt "..." --output OUT
python -m cyllama.sd video --model M --prompt "..." --output frames/
# Upscale image
python -m cyllama.sd upscale \
--model models/esrgan-x4.bin \
--input image.png \
--output image_4x.png
# Convert model
python -m cyllama.sd convert \
--input sd-v1-5.safetensors \
--output sd-v1-5-q4_0.gguf \
--type q4_0
# Show system info
python -m cyllama.sd info
Supported Models¶
-
SD 1.x/2.x: Standard Stable Diffusion models
-
SDXL/SDXL Turbo: Stable Diffusion XL (use cfg_scale=1.0, steps=1-4 for Turbo)
-
SD3/SD3.5: Stable Diffusion 3.x
-
FLUX: FLUX.1 models (dev, schnell)
-
Wan/CogVideoX: Video generation models (use
generate_video()) -
LoRA: Low-rank adaptation files
-
ControlNet: Conditional generation with control images
-
ESRGAN: Image upscaling models
Error Handling¶
All cyllama functions raise appropriate Python exceptions:
from cyllama import complete, LLM
try:
response = complete("Hello", model_path="nonexistent.gguf")
except FileNotFoundError:
print("Model file not found")
except RuntimeError as e:
print(f"Runtime error: {e}")
except Exception as e:
print(f"Unexpected error: {e}")
# LLM with error handling
try:
gen = LLM("models/llama.gguf")
response = gen("What is Python?")
except Exception as e:
print(f"Generation failed: {e}")
Type Hints¶
All functions include comprehensive type hints for IDE support:
from typing import List, Dict, Optional, Iterator, Callable, Tuple
from cyllama import (
complete, # str | Iterator[str]
chat, # str | Iterator[str]
LLM, # class
GenerationConfig, # @dataclass
)
Performance Tips¶
1. Model Reuse¶
# BAD: Reloads model each time (slow)
for prompt in prompts:
response = complete(prompt, model_path="model.gguf")
# GOOD: Reuses loaded model (fast)
gen = LLM("model.gguf")
for prompt in prompts:
response = gen(prompt)
2. Batch Processing¶
from cyllama import batch_generate, GenerationConfig
# BAD: Sequential processing
responses = [generate(p, model_path="model.gguf") for p in prompts]
# GOOD: Parallel batch processing (3-10x faster)
prompts = ["What is 2+2?", "What is 3+3?", "What is 4+4?"]
responses = batch_generate(
prompts,
model_path="model.gguf",
n_seq_max=8, # Max parallel sequences
config=GenerationConfig(max_tokens=50, temperature=0.7)
)
3. GPU Offloading¶
# Estimate optimal layers
from cyllama import estimate_gpu_layers
estimate = estimate_gpu_layers("model.gguf", gpu_memory_mb=8000)
# Use recommended settings
config = GenerationConfig(n_gpu_layers=estimate.layers)
gen = LLM("model.gguf", config=config)
4. Context Sizing¶
# Auto-size context (recommended)
config = GenerationConfig(n_ctx=None, max_tokens=200)
# Manual sizing (for control)
config = GenerationConfig(n_ctx=2048, max_tokens=200)
5. Streaming for Long Outputs¶
# Non-streaming: waits for complete response
response = complete("Write a long essay", model_path="model.gguf", max_tokens=2000)
# Streaming: see output as it generates
for chunk in complete("Write a long essay", model_path="model.gguf",
max_tokens=2000, stream=True):
print(chunk, end="", flush=True)
Version Compatibility¶
-
Python: >=3.12 (abi3 wheels; tested on 3.13). For 3.10/3.11 use
cyllama<0.3.0or build from source. -
llama.cpp: see
CHANGELOG.mdfor the pinned revision -
Platform: macOS, Linux, Windows
See Also¶
-
User Guide - Comprehensive usage guide
-
Cookbook - Practical recipes and patterns
-
Changelog - Release history
See pyproject.toml for the current cyllama version.