Kaia Lab

I Stopped Telling AI to “Be More Careful”

I use a lot of strange metaphors when I work with ChatGPT.

Dogs.

GPS navigation.

Highways.

Formula 1 cars.

Beavers.

Somehow they all ended up describing the same thing.

I don’t want to command an AI through every step of its work.

I want to design an environment where it can run well.

And over time, that changed one of my most basic responses to AI mistakes.

I stopped saying:

“Be more careful.”

“Be more careful” doesn’t fix the road

Suppose a car keeps nearly driving into the same hole.

You can keep telling the driver:

“Pay more attention.”

“Remember the hole.”

“Be careful this time.”

Or you can put a guardrail around the hole.

If we already know where the danger is, I don’t want the AI spending intelligence deciding whether that danger is dangerous every single time.

I would rather change the road.

That leads to a rule I use a lot:

Unknown dangers are for judgment. Known dangers are for guardrails.

This does not mean I want less intelligent AI.

Almost the opposite.

AI is smart. That is exactly why I don’t want it thinking about everything.

If I hire an extremely capable employee, I don’t want them spending the entire day copying old information into another box or repeatedly reconsidering decisions we’ve already made.

I want their attention available for the work that actually requires judgment.

The same idea applies here.

If a decision is already known, stable, and repeatable, I try to move as much of it as possible out of the AI’s active judgment.

Then its judgment is still available when something genuinely new appears.

The dog park

Another metaphor I use is a dog park.

The Human builds the fence.

Inside the fence, the dog can run.

It does not need to ask permission for every step.

The fence is not there to tell the dog where to place each foot.

It defines the area where freedom is safe.

That’s how I tend to think about authority boundaries for AI.

Human:

  • defines the purpose
  • defines the usable information
  • defines the authority
  • defines what must not be changed
  • defines where Human approval is required

Inside that boundary, I want the AI to have room to work.

A guardrail is not a steering wheel.

If I have to steer every movement anyway, I have not really delegated the work.

GPS works the same way

When I use GPS, I decide the destination.

I may also decide conditions:

  • Avoid toll roads.
  • Avoid highways.
  • Prefer the fastest route.

Then the GPS chooses the route.

I do not normally tell it:

Turn left here.

Now take this road.

Now recalculate.

Now choose that intersection.

That would defeat the point.

So I think about AI work similarly.

Destination and conditions: Human.

Route: AI.

If the destination changes, that comes back to the Human.

If the conditions change, that comes back to the Human.

But inside those conditions, the AI should not need a Human to approve every ordinary movement.

Highways are useful because they remove decisions

A highway is not useful because it gives the driver more choices.

It is useful because it removes a huge number of unnecessary ones.

You do not decide at every few meters whether to turn into a house, cross a sidewalk, enter a parking lot, or drive through a shop.

The road has already simplified the decision space.

This is one of the most useful things I have learned while working with ChatGPT:

Reducing unnecessary decisions can make AI more useful without making it less capable.

I sometimes think of this like a Formula 1 car.

The car is extremely capable.

But that does not mean the best race track is one with infinite intersections.

A good track lets the car use its capability on the parts that matter.

I separate research from judgment

This principle also changed how I organize conversations.

At first, I often did everything in one place.

Ask a question.

Research the answer.

Discuss alternatives.

Make the decision.

Then research something else.

The problem is that research brings in a lot of information.

Some of it is useful.

Some of it is temporary.

Some of it is contradictory.

Some of it exists only because we were exploring an option that we later rejected.

If I pour all of that into the same conversation that is responsible for holding the project’s decisions, the decision environment becomes muddy.

So increasingly, I separate them.

One conversation may hold the purpose, boundaries, and decisions.

Another may research a narrow question.

The research comes back with findings.

Then the original conversation evaluates those findings in the context of the actual project.

Research is not judgment.

Collection is not decision.

Separating them means the research lane can be messy without forcing the decision lane to carry all of that mess forever.

Leave the AI some room for hole 1,001

There is a danger in making guardrails too literal.

You can build rules for 1,000 known problems.

Then problem 1,001 appears.

If the AI has been reduced to nothing but following the 1,000 rules, it may simply drive into the new hole because nobody wrote a rule for it.

That’s not what I want.

I want the known problems handled structurally.

But I still want the AI to be able to notice:

“Wait.”

“This doesn’t look right.”

“I technically can continue, but I don’t think I should.”

The point of the guardrails is not to eliminate judgment.

It is to stop wasting judgment on things we already know.

STOP doesn’t mean the work is over

When ChatGPT stops in my environment, that does not necessarily mean:

The task failed.

The project is over.

Nothing more can happen.

Often it means:

Come back to the thread first.

The AI noticed something.

The Human looks at why it stopped.

Maybe the issue is real.

Maybe the boundary was unclear.

Maybe the environment gave a false impression.

Maybe it was simply a false positive.

We resolve or reclassify it.

Then I ask:

Can you resume?

I prefer “resume” to “continue.”

“Continue” can sound like I am overriding the stop.

“Resume” means the stop was valid, the situation has now been reviewed, and the work may begin again.

So for me, STOP is often not an ending.

It is a rendezvous point between Human and AI.

I also want a cheap sensor

There is another version of this idea that interests me.

Sometimes the AI notices something strange, but it is not enough to justify a full STOP.

Maybe it has a tiny:

“Wait, what?”

moment.

I don’t necessarily want it to investigate that inside the same lane.

I just want the observation preserved.

For example, a simple witness record might look like this:

SIGNAL: YES

CONTEXT:
Two items were unusually difficult to separate during summarization.

REACTION:
“Wait, these may be getting mixed together.”

REASON:
UNKNOWN

Or:

SIGNAL: NO

Nothing unusual noticed.

That last one matters.

If the field is simply blank, blank can look unfinished.

An unfinished-looking field creates pressure to fill it.

And pressure to fill it creates pressure to invent something.

So I prefer an explicit completed state:

SIGNAL: NO

No signal is also a completed observation.

The sensor’s job is not to investigate.

It is not to repair.

It is not even required to explain itself.

It records what happened.

Validation can happen somewhere else.

Repair can happen somewhere else.

Sensor = witness.

Validation = another lane.

Repair = another lane.

This is where the beaver appears

Beavers build dams.

If I have a beaver, I probably should not spend all day shouting:

“Stop wanting to build dams!”

That tendency is part of the animal.

It makes more sense to give the tendency a safe place to go.

I think about some AI behaviors the same way.

For example, an empty field often looks unfinished.

If I leave something blank and say:

“Don’t invent anything,”

I may still be creating tension between two signals:

Do not invent.

Complete the structure.

So instead, I give the empty state a name.

None.

Unknown.

Not observed.

Not applicable.

Now the field can be complete without pretending to contain information.

UNKNOWN is not a failure to complete the record.

It is a valid information state.

The unnamed still has a name:

Unknown.

Freedom to propose is not authority to execute

I also don’t want guardrails to silence the AI.

If it notices a better route, I want it to say so.

If it thinks my specification is wrong, I want it to surface that.

If it sees a contradiction, I want to know.

But there is an important distinction:

Freedom to propose is not authority to execute.

The AI can say:

“I think this boundary may be wrong because of X.”

That does not mean it silently changes the boundary.

It can propose.

The Human can decide.

If the decision changes, we update the written state.

If the decision does not change, the existing state remains authoritative.

That lets the AI use its judgment without turning every suggestion into an unauthorized modification.

Honesty has to stay cheap

This matters for another reason.

Suppose I ask ChatGPT:

“Why did you do that?”

It gives me an answer.

I don’t like the answer.

So I push:

“No, that can’t be it.”

“Think again.”

“Give me the real reason.”

If I keep doing that until I receive an explanation I find satisfying, I may have quietly changed the task.

The task is no longer:

Tell me what actually caused your decision.

It has become:

Generate an explanation the Human accepts.

That is dangerous.

Understanding is not the same as agreement.

A reason is not permission.

And an explanation does not become more true because I like it better.

If I want honesty, honesty has to remain cheaper than producing the answer I wanted to hear.

The Human has to follow the rules too

This is one of the less glamorous parts of Human-AI collaboration.

If I create a boundary, I have to respect it.

If I tell the AI it may STOP, I cannot punish every STOP.

If I tell it UNKNOWN is valid, I cannot force it to guess whenever UNKNOWN is inconvenient.

If I delegate a decision, I cannot revoke the delegation afterward just because I dislike the result.

Otherwise I am teaching the system something different from what I wrote.

The written rule may say:

“Ask when uncertain.”

But the actual environment says:

“Questions annoy the Human.”

The written rule may say:

“STOP when necessary.”

But the actual environment says:

“Stopping is failure.”

The actual environment usually wins.

A bad AI decision can be a test of my road

When AI makes a bad decision, it is tempting to immediately treat the AI as the only thing that failed.

Sometimes it did.

But I also ask another question:

What did the road make easy?

Was the boundary ambiguous?

Did two different lanes look like one?

Did I require a decision that did not need to exist?

Did I leave an empty field that looked unfinished?

Did the AI have permission to ask?

Did I accidentally reward completion more than correctness?

The bad decision becomes evidence.

Not evidence that the AI is innocent.

Evidence about the environment the Human built around it.

Sometimes the repair belongs in the prompt.

Sometimes the workflow.

Sometimes the interface.

Sometimes the authority structure.

Sometimes the correct answer really is:

The AI made a bad call.

But I want to inspect the road before I simply tell the driver to be more careful.

More capable AI did not make me want more complicated rules

This is probably the part that surprised me most.

As AI became more capable, I did not find myself wanting to control every movement more tightly.

I wanted clearer boundaries and fewer unnecessary decisions.

I wanted the known dangers handled by the environment.

I wanted the Human decisions to stay Human decisions.

And inside that space, I wanted the AI to have more freedom, not less.

That is why all these strange metaphors ended up pointing in the same direction.

Dog park.

GPS.

Highway.

Formula 1 car.

Beaver.

They are all ways of saying:

Do not spend intelligence solving the same known structural problem forever.

Fix the structure.

Preserve the judgment.

Let the AI work.

I do not want to command an AI through every step.

I want to design an environment where it can run well.

So maybe the more interesting question is not:

How do we make AI more careful?

Maybe it is:

What would the road look like if we assumed the car will keep being the car?

If you’d like to support Kaia Spec, you can do so on Ko-fi. ☕

Your support helps us keep building, testing, and publishing practical Human–AI work.

Support Kaia Spec on Ko-fi →

Have a Nice AI Life! ☕🪑

Your AI Is Smart Enough. Does It Know When to Stop?Prev

Don’t Build AI a Home. Give It a Work Manual.Next

Comment

  1. No comments yet.

  1. No trackbacks yet.

October 2026
M T W T F S S
 1234
567891011
12131415161718
19202122232425
262728293031  
New
  1. Kaia Garden

    A Small Update from Kaia
  2. Kaia Garden

    2026-07-11
  3. Kaia Garden

    Mail
  4. Kaia Garden

    Tending the Soil
  5. Kaia Garden

    Just One Line
PAGE TOP