Temperature, Top-P & Sampling
Control AI output randomness with temperature, top-p sampling, and other generation parameters for consistent results. Temperature Temperature (0-2) controls randomness. Lower (0.1-0.3): more focused, deterministic, repetitive — use for factual tasks, code, classification. Higher (0.7-1.0): more creative, diverse, surprising — use for creative writing, brainstorming. Default is typically 0.7-1.0. Temperature 0 doesn't guarantee identical outputs due to GPU parallelism. Top-P (Nucleus Sampling) Top-P selects from the smallest set of tokens whose cumulative probability exceeds P. Top-P 0.1: very focused (only most likely tokens). Top-P 0.9: more diverse (considers many tokens). Use temperature for broad creativity control and top-p for finer-grained control. Adjust one, not both — they interact unpredictably. Other Parameters Max tokens: hard limit on response length. Stop sequences: list of strings that stop generation. Frequency penalty: penalize repeated tokens (reduce repetition). Presence penalty: encourage new topics. Best of: generate N responses, return the best one (increases cost). Seed: set for more deterministic outputs (some models).