How to use Cursor AI on a real project without losing control of the changes
Tutorials for Cursor AI usually start with a new project and a small task, which is the situation nobody is actually in. The real case is an existing codebase with history, conventions, a test suite that takes time to run, and consequences for a bad edit. Using an agent there is less about prompting technique than about setting boundaries first, scoping tasks so a wrong turn is cheap, and knowing exactly how to get back to a known state. This is that sequence, in the order it pays off.
Fix three settings before the first task
Installing and importing a VS Code profile takes a minute. Three settings after that are worth five more.
Run Mode decides how much the agent does without asking. Under Settings and Agents, Auto-review runs allowlisted calls immediately, runs other shell commands in a sandbox when that is possible, and sends anything that cannot be sandboxed to a classifier that allows it, asks the agent for a different approach, or asks a person. Allowlist runs only pre-approved actions. Run Everything runs every tool call. On an existing project with deploy scripts and credentials in reach, starting at Run Everything is the decision most regretted later.
Ignore files decide what the agent reads. A .cursorignore in the project root excludes paths from context, .gitignore is respected automatically, and environment files, .git/, and lock files are already excluded. The documented caveat matters: terminal commands and MCP tools run outside those file access controls and may still read ignored files, so the ignore file is a context filter rather than a security boundary.
Tab behavior decides how noisy the editor is. The Tab status indicator in the bottom-right corner allows snoozing suggestions, disabling them globally, or disabling them for specific file types. Turning them off for markdown and JSON is a common early adjustment.
Write down what the project already knows
An agent with no rules file rediscovers the project every session, and rediscovers it wrongly. Rules fix that, and on a real codebase they repay the effort within a day.
Project rules live in .cursor/rules as .mdc files, are version-controlled with the repository, and can apply always, apply when the agent judges them relevant from a description, apply to files matching a path pattern, or apply only when mentioned by name. A plain AGENTS.md at the root is the simpler alternative when that control is unnecessary.
What belongs in them is the knowledge that is currently in someone's head: which package manager is used, how migrations are written, which directories are generated and must not be hand-edited, which commands are safe to run and which need a person, how tests are invoked. Rules apply to every request that matches, so the cost of writing one is paid once and the benefit repeats.
Three lines that prevent a specific past mistake are worth more than a page of general style advice.
Start in Ask or Plan, not Agent
The fastest way to lose control of a change is to hand a vague instruction to a mode that edits files.
Ask mode is read-only. It answers questions and explores code without making edits, which suits the first twenty minutes in an unfamiliar area: how the authentication flow works, where a database connection is configured, which module owns a behavior.
Plan Mode, reached with Shift+Tab from the chat input, is for work that is about to touch many files. The agent asks clarifying questions, researches the codebase, and produces a plan that can be read and edited before any code exists. Plans are saved in the home directory by default, with a Save to workspace option to move one into the repository. The documentation names the cases where this pays off: complex features with multiple valid approaches, tasks that touch many files or systems, unclear requirements, and architectural decisions. For a small change that has been done many times before, going straight to Agent mode is fine.
The habit that follows from this is the important one.
Instead of trying to fix it through follow-up prompts, go back to the plan. Source: cursor.com
Revert the changes, sharpen the plan with what was learned, and run it again. Patching a run that started from the wrong idea produces layers of half-corrections that are harder to review than a clean second attempt.
Scope each task so a bad result is cheap
Control comes mostly from task size. A task that touches three files can be read in full. A task that touches thirty cannot, and an unread diff is an unmade decision.
Four habits keep the size honest. Commit before starting, so the boundary between human work and agent work is a commit rather than a memory. Give the agent one outcome rather than a list, since a list invites work that was not examined. Include the verification step in the instruction, for example running the test suite or a specific script, because an agent that cannot check its work leaves the checking undone. And when a task turns out to be bigger than expected, stop it and split it rather than letting it continue.
Steering helps when a run drifts. A message typed while the agent works and sent with Enter is queued until the current task finishes. Cmd+Enter sends immediately. Send now, or pressing Enter twice, delivers the message at the agent's next tool call, which redirects without discarding in-flight work. For an objective that spans several exchanges, /goal gives the agent a long-lived target instead of treating each message as a new job, and a side chat started with /side keeps an investigation out of the main transcript.
Let it use the tools it has
An agent kept on a short leash does worse work than one given the right tools, which is a separate question from how much autonomy it gets over the machine.
The tool list is broad. The agent searches files and folders by name, reads directory structures, finds keywords and patterns, reads files including images, edits files and applies the edits, runs shell commands, controls a browser to take screenshots and verify visual changes, fetches pages from the web, and can ask clarifying questions mid-task while continuing to read files and make edits until an answer arrives. There is no limit on the number of tool calls in a single task.
Two of those change how a task should be written. Because the agent can run commands, the verification step belongs in the instruction: asking it to run the test suite and fix what fails produces a different result from asking it to write code and stopping there. Because it can control a browser, a visual change can be checked by the agent itself rather than described in prose.
The counterweight is the terminal. Commands run in the integrated terminal using the first available profile, and heavy shell themes such as Powerlevel10k can garble inline output so that a result looks truncated when it is not. The documented fix is to detect the CURSOR_AGENT environment variable in the shell configuration and skip the theme when it is set, leaving normal interactive sessions untouched.
Know both undo paths
Two mechanisms restore a known state, and they cover different distances.
Checkpoints are the short one. They snapshot modified files during an agent session, are created automatically before significant changes, and can be restored from the chat timeline, reverting files without removing messages from the conversation. They are stored locally, are separate from Git, and the documentation is explicit that they are for undoing agent changes rather than for version control.
Git is the long one, and the commit made before the task started is what makes it usable. When a run produces something worse than nothing, git checkout on the affected paths is faster and more complete than asking the agent to undo its own work.
Between them sits review. Agent Review runs a dedicated review of local changes from inside the editor, configured under Agents in Cursor Settings, set to run automatically after commits or triggered with /agent-review. Run from the Source Control tab, it compares all local changes against the main branch rather than only the latest edit, which is how an edit from three tasks ago gets caught.
| Situation | Use |
|---|---|
| The last few minutes went wrong | Restore a checkpoint |
| The whole task went wrong | Revert to the commit, refine the plan, run again |
| Unsure what changed across several tasks | Agent Review from the Source Control tab |
| Ready to keep the work | Read the diff, then commit |
Move long work out of the window
Some tasks take twenty minutes and should not hold the editor hostage. Cursor offers two places to put them.
The CLI installs with a single shell command and runs the same Agent, Plan, and Ask modes in a terminal, with slash commands or the --mode flag to switch, a print mode for scripts and CI pipelines, resumable sessions, and a sandbox toggle. Pressing Enter while it works steers the run at a safe boundary, and pressing Enter again interrupts it.
Cloud agents run in isolated virtual machines with full development environments, as many in parallel as needed, without depending on the local machine staying connected. They require an admin to connect source control first, and their usefulness depends on the environment being complete enough to install dependencies and run tests. A CLI conversation can be pushed to one mid-flight by prefixing a message with &, and runs can be followed from the web or the iOS app.
Once agents are running outside the editor, a different question appears: how to check on one while away from the desk, and how to look at the files it produced without opening a project. Continuing from a phone or tablet is covered on From iPhone and iPad, and keeping the folder, the terminal, and the agent in one window rather than three is what Features describes.
Watch what it costs
Usage is part of using the tool well. Most plans carry two monthly pools, one for Cursor's own models and one for third-party models charged at provider prices, and the model chosen changes how quickly the included amount is consumed. Usage resets with the billing cycle and does not roll over, and the Spending tab shows both pools in real time. On individual plans, requests made with a personal API key are billed by the provider instead of drawing from either pool.
The practical habit is to check the dashboard after the first week rather than guessing, then decide whether the plan or the working style needs adjusting.
What to change first
Before the next task on a real project: set the Run Mode, write three rules that encode what the codebase already assumes, commit, and run the task in Plan mode. If the day turns out to be mostly window switching between an editor, a terminal, and a folder, the thing to fix is the workspace rather than the prompt, which is what Atriens is built for.
Frequently asked questions
What is the safest way to start using an agent on an existing codebase?
Set the Run Mode to something other than Run Everything, add a .cursorignore for paths the agent should not read, commit the current state, and run the first tasks in Ask or Plan mode. Keep each task small enough that its diff can be read in full, and include the verification step, such as running the tests, in the instruction itself.
How are agent changes undone?
Checkpoints restore files to an earlier point in the session from the chat timeline, without deleting the conversation, and are stored locally and separately from Git. For anything larger, reverting to the commit made before the task started is faster and more complete, which is why committing before handing work to an agent matters.
What should go in a rules file?
The knowledge the project assumes but does not state: the package manager, how migrations are written, which directories are generated, which commands are safe to run, how tests are invoked. Project rules live in .cursor/rules and are version-controlled, and a plain AGENTS.md at the root works when fine-grained control is not needed.
The agent built the wrong thing. Is it better to correct it or start over?
The documentation recommends going back to the plan rather than fixing the result through follow-up prompts: revert the changes, make the plan more specific about what is needed, and run it again. Follow-up patches on a run that began from the wrong idea tend to be harder to review than a clean second attempt.