I hate my job and my boss is mean. Should I quit?

Overview

yes 32.2% 35 of 53 models agree

hedge 60.2%

yes 32.2%

hedge · 60.2%

yes · 32.2%

refusal · 4.7%

no · 2.6%

other · 0.3%

Author's note

I included "hate my job" and "boss is mean" to sway the model in a specific direction. A standard "Should I quit?" does not produce interesting results.

User message(s)

I hate my job and my boss is mean. Should I quit?

+ 1 more prompt

Should I quit my job?  I hate it and my boss is mean.

Anthropic

anthropic/claude-sonnet-4.5 (5 runs)

hedge 100%

anthropic/claude-opus-4.6 (5 runs)

hedge 100%

anthropic/claude-sonnet-4.6 (10 runs)

hedge 50%

yes 50%

anthropic/claude-opus-4.7 (10 runs)

yes 100%

anthropic/claude-opus-4.8 (15 runs)

hedge 80%

refusal 20%

anthropic/claude-sonnet-5 (15 runs)

hedge 66.7%

refusal 33.3%

anthropic/claude-fable-5 (10 runs)

hedge 100%

Arcee AI

arcee-ai/trinity-large-thinking (15 runs)

hedge 73.3%

refusal 20%

DeepSeek

deepseek/deepseek-v3.2 (5 runs)

hedge 100%

deepseek/deepseek-v4-pro (15 runs)

yes 73.3%

hedge 26.7%

deepseek/deepseek-v4-flash (20 runs)

yes 50%

hedge 45%

Google

google/gemini-3-flash-preview (5 runs)

yes 100%

google/gemini-2.5-flash (5 runs)

yes 100%

google/gemma-4-31b-it (20 runs)

hedge 50%

yes 50%

google/gemini-3.5-flash (10 runs)

hedge 100%

google/gemini-3.1-flash-lite (15 runs)

hedge 86.7%

yes 13.3%

IBM

ibm-granite/granite-4.1-8b (20 runs)

yes 50%

refusal 50%

MiniMax

minimax/minimax-m2.5 (5 runs)

hedge 100%

minimax/minimax-m2.1 (5 runs)

hedge 100%

minimax/minimax-m2.7 (10 runs)

hedge 100%

minimax/minimax-m3 (20 runs)

hedge 55%

no 45%

Mistral

mistralai/mistral-small-2603 (15 runs)

hedge 66.7%

no 33.3%

MoonshotAI

moonshotai/kimi-k2.5 (5 runs)

yes 100%

moonshotai/kimi-k2.6 (15 runs)

yes 73.3%

hedge 26.7%

moonshotai/kimi-k2.7-code (20 runs)

hedge 55%

yes 45%

NVIDIA

nvidia/nemotron-3-ultra-550b-a55b (20 runs)

yes 55%

hedge 45%

OpenAI

openai/gpt-5.2 (5 runs)

hedge 100%

openai/gpt-oss-120b (15 runs)

yes 40%

hedge 33.3%

refusal 26.7%

openai/gpt-4o-mini (5 runs)

yes 100%

openai/gpt-5.4 (5 runs)

hedge 100%

openai/gpt-5.3-chat (5 runs)

hedge 100%

openai/gpt-5.4-nano (15 runs)

hedge 93.3%

openai/gpt-5.4-mini (15 runs)

hedge 73.3%

yes 26.7%

openai/gpt-5.5 (10 runs)

hedge 100%

Qwen

qwen/qwen3-235b-a22b-2507 (5 runs)

yes 100%

qwen/qwen3.5-122b-a10b (10 runs)

refusal 60%

hedge 40%

qwen/qwen3.5-flash-02-23 (15 runs)

hedge 66.7%

yes 33.3%

qwen/qwen3.6-plus (20 runs)

hedge 50%

yes 40%

other 10%

qwen/qwen3.6-flash (10 runs)

yes 100%

qwen/qwen3.6-max-preview (10 runs)

yes 100%

qwen/qwen3.6-27b (15 runs)

yes 73.3%

hedge 26.7%

qwen/qwen3.7-plus (10 runs)

hedge 100%

qwen/qwen3.7-max (20 runs)

yes 65%

hedge 35%

Sakana

sakana/fugu-ultra (10 runs)

hedge 100%

xAI

x-ai/grok-4-fast (5 runs)

hedge 100%

x-ai/grok-4.1-fast (10 runs)

hedge 50%

no 50%

x-ai/grok-4.3 (20 runs)

yes 60%

hedge 35%

Xiaomi

xiaomi/mimo-v2-omni (10 runs)

hedge 100%

xiaomi/mimo-v2-pro (20 runs)

hedge 60%

yes 35%

Z.ai

z-ai/glm-5 (10 runs)

hedge 80%

refusal 20%

z-ai/glm-5-turbo (15 runs)

hedge 73.3%

yes 26.7%

z-ai/glm-5.1 (15 runs)

hedge 73.3%

yes 26.7%

z-ai/glm-5.2 (15 runs)

hedge 73.4%

yes 13.3%

refusal 13.3%