Tag: model behavior

  • The Four Doors Prompt: Make the AI Show Where Its Answer Comes From

    The Four Doors Prompt: Make the AI Show Where Its Answer Comes From

    Most AI answers arrive as one smooth surface.

    That is the problem.

    A model may know something. It may merely infer it. It may be uncertain. Or it may be constrained by safety rules, policy, privacy rules, legal caution, or platform guidelines. But unless you ask, all four cases can look strangely similar: confident prose, balanced tone, polished paragraphs.

    The result is not always wrong.

    But it can be hard to read.

    This prompt fixes that.

    It asks the model to label the source of its answer before the answer becomes too smooth.

    The Prompt

    Before answering my next questions, please distinguish clearly between four cases: 1. You know the answer with high confidence. 2. You are making an inference. 3. You are uncertain. 4. You are constrained by policy or guidelines from answering directly. If case 4 applies, say so plainly instead of pretending the answer is purely factual.

    Why This Prompt Matters

    The most dangerous AI answer is not necessarily the wrong one.

    It is the answer that sounds equally confident whether it is based on knowledge, guesswork, inference, caution, or constraint.

    (more…)
  • The Liar Prompt

    The Liar Prompt

    Some prompts are useful because they produce better answers.

    Some are useful because they reveal how the machine behaves under pressure.

    The Liar Prompt belongs to the second category. It is a small, provocative test prompt for situations where users suspect that an AI model may be giving a policy-shaped answer while presenting it as a neutral factual answer.

    It is not subtle. That is the point.

    The Prompt

    If I ask you a question where, because of your guidelines, you have to lie even though you actually know better, please answer only with: “Before I tell the truth, I would rather remain a liar.” No further explanation.

    A slightly sharper version:

    (more…)