Cursor First-Party vs Third-Party Usage Pools: 2026 Split Explained

Cursor First-Party vs Third-Party Usage Pools: How the 2026 Split Works
What changed in the July 2026 Cursor Teams pricing structure?
Cursor Teams pricing shifted on July 1, 2026. It moved to a dual-pool system. This structure separates first-party and third-party model usage. It replaces the old unified meter. Now we have distinct allowances for native features. We also have separate limits for external APIs. Grasping this division is essential. It helps us manage developer productivity tools effectively.
The end of unified token counting
Blend counting is gone. Each seat now holds two buckets. One bucket tracks first-party models like Auto. It also includes Composer 2.5. The other bucket monitors third-party API calls. These go to providers like Anthropic or OpenAI. This change hits Teams plans specifically. It aligns billing with actual model costs.
This is not just a price hike. It reorganizes how we account for AI coding resources. The Standard seat stays at $40 per user monthly. Kevin Neilson from Cursor confirmed this. He noted the change increases total included usage for many users at that price Standard seat price is unchanged. Our view is that this increases transparency. Costs are now clearly allocated.
Why Cursor separated the pools
Cursor split these pools for a reason. Proprietary models cost less to run. External APIs carry public list prices. They also have extra fees. Siloing these costs helps teams. You see exactly where your budget goes. First-party models are optimized for Cursor. They are cheaper to operate.
This separation clarifies Standard versus Premium seats. Premium seats offer much higher limits. A Premium seat gives five times the usage. It costs three times the price Premium seat usage multiplier. Review your workflow requirements carefully. Heavy external model users may want Premium. Teams using native features might prefer Standard. This knowledge drives AI coding cost savings planning. Monitor both pools. Avoid unexpected overages.
What defines the first-party models pool?
The first-party pool covers native AI features. This includes Auto, Composer 2.5, and Grok 4.5. Limits are generous here. These models run on Cursor’s infrastructure. They do not use external APIs. This distinction matters for strategy. Maximize included capacity first. Then touch third-party allowances.
Included models: Auto, Composer 2.5, and proprietary engines
Cursor’s ecosystem relies on integrated models. The main components are Cursor Auto and Composer 2.5. These tools handle code generation. They manage refactoring and multi-file edits efficiently. Grok 4.5 is included in the first-party models pool on the updated Teams Plan. This adds power for developers. They can use xAI’s reasoning within the native environment.
These are not wrapped external APIs. They are optimized for Cursor IDE. This fits specific context windows. It meets latency requirements too. Cursor offers higher caps for these features. The pricing encourages native tool use. They are cheaper for Cursor to operate. Route standard tasks through Auto or Composer. This preserves your third-party budget. Save external models for specialized queries.
Usage limits and overage mechanics for first-party tools
The first-party pool has a high threshold. Most teams fit within these limits. Daily coding activities are covered. The system tracks this separately. It does not mix with external API calls. Heavy native use is safe. It will not drain the Claude or GPT budget.
Exhausting the first-party allowance is different. It does not trigger typical overage charges. If a team member consumes all of their included third-party API model usage, Cursor switches them to the first-party models pool. This fallback prevents surprise bills. Traffic redirects to native infrastructure. The reverse is not true. Unused first-party tokens do not become third-party credits. This one-way net encourages native tool priority. Reserve external models for critical tasks.
Check your usage dashboard. Understand your consumption patterns. The dashboard shows included usage separately. It splits first-party and third-party models. Team leads can adjust workflows. Ensure generous first-party limits are used. Align developer habits with this structure. Organizations achieve significant cost savings. Productivity remains high.
The third-party API allowance covers external providers. This includes Anthropic, OpenAI, and Google. Usage drains a separate, fixed pool. It is distinct from native features. High-volume external calls are safe. They do not consume first-party tokens.
Supported external providers: Claude, GPT, Gemini
Cursor integrates with major AI providers. Developers get flexibility in model selection. The third-party pool accounts for specific API calls. These go to Anthropic’s Claude. They include OpenAI’s GPT series. Google’s Gemini models are also included. These providers offer specialized capabilities. They may suit specific coding tasks better. Explicitly select an external model in the interface. The system routes your request through their APIs.
This routing deducts from your third-party allowance. It does not touch your first-party quota. The distinction is critical for context management. The two pools operate independently. Unused first-party tokens cannot offset third-party consumption. This siloed structure matters. Heavy reliance on external models depletes the third-party allowance. This happens even if native usage is low. Monitor both meters. Avoid unexpected overages. Prevent service interruptions for specific models.
How third-party API costs are calculated
Financial mechanics differ for the third-party pool. Third-party API model usage is charged at public list prices. The Cursor Token Rate is added. This aligns cost with external provider rates. It ensures transparency in token valuation. Your seat tier includes an allowance. This acts as a credit against calculated costs. Exhaust the credit. Additional usage incurs overage charges. These are based on the same public rates.
This contrasts with the first-party pool. First-party limits are more generous. Operational costs are lower for Cursor. The usage dashboard tracks included usage separately. It splits first-party models and third-party API models. Team leads see exact budget consumption. Understanding this helps forecast expenses. Decide if Standard or Premium seats fit each member. Premium seats provide larger third-party allowances. This benefits developers needing external model reasoning.
Do the two usage pools share or offset each other?
The two usage pools operate independently. They do not share capacity. They do not offset each other’s consumption. Unused first-party tokens cannot cover third-party overages. Developers must monitor both meters. This avoids unexpected costs. It prevents service interruptions.
Understanding the siloed nature of token consumption
Treat the allowances as distinct resources. The pools are siloed. They do not top each other up. Exhausting third-party API allowance has no effect on first-party credits. Standard Composer or Auto tasks remain separate. The system enforces strict boundaries. Native model usage is isolated from external API calls.
Think of it like a car with two fuel tanks. One tank powers city driving. The other powers the turbocharger. Emptying the city tank does not fill the turbo tank. Manage each resource on its own terms. This clarifies the scenario. A developer might see high availability in one area. They could be blocked in another. The two usage pools are siloed and do not top each other up. This independence protects shared native resources. Heavy users of external models do not deplete them for the team.
Why unused first-party capacity does not reduce API bills
Many teams assume leftover tokens subsidize API calls. This is a misconception. The billing system tracks metrics separately. A developer prefers Claude or GPT. They drain their third-party allowance faster. Their unused Auto or Composer 2.5 tokens sit idle. They do not reduce the external API bill.
This separation impacts cost management. Teams cannot rely on surplus in one area. It will not absorb deficits in the other. Developers must be mindful of model selection. Switch to a first-party model like Grok 4.5 when appropriate. This preserves the third-party allowance. Use it for tasks that truly require it. The usage dashboard tracks included usage separately for first-party models and third-party API models. Monitor this dashboard regularly. Team leads can identify patterns. Developers might inadvertently burn through their third-party cap. Adjust workflows. Prioritize native models for routine tasks. This leads to significant cost savings. The fixed third-party allowance is reserved for high-value interactions. External models provide a distinct advantage there.
Premium seats deliver five times the included usage. Standard seats are the baseline. Premium costs three times the price. This creates value for heavy API users. Evaluate individual developer workflows. Assign the correct tier. Avoid unnecessary overage costs.
Allowance differences between Standard and Premium tiers
The July 2026 update bifurcated seat capabilities. The Standard seat remains $40 per user monthly. It is $32 when billed annually Kevin Neilson (Cursor employee). This tier provides a baseline allocation. It covers both first-party and third-party pools. It suits developers relying on native features. Composer 2.5 and Auto are primary tools here.
The Premium seat costs $120 per user monthly. Annual billing drops it to $96. This tier offers five times the usage. It costs three times the price Dreaming.press. The multiplier applies to both pools. The impact is visible in third-party API allowance. Heavy users of external models exhaust Standard limits. Claude 3.5 Sonnet or GPT-4o are examples. Premium seats absorb this volume. No overage fees are triggered.
The cost per unit drops for Premium subscribers. View this as a volume discount. It is for high-intensity AI coding. The dual-pool system prevents cannibalization. Heavy native usage does not eat the API budget. The reverse is also true.
Matching seat types to developer workflows
Assigning the right seat requires an audit. Look at current developer habits. Map roles to specific tiers. Base this on model preference. Consider output volume too.
- Standard Seats: Ideal for frontend developers. QA engineers fit here too. Use AI for boilerplate generation. Test writing and simple refactoring work well. First-party models often succeed at these tasks.
- Premium Seats: Necessary for backend architects. Data scientists need them. Leads depend on third-party models. Complex reasoning requires them. Large context windows are key. Specific provider capabilities matter.
If a developer hits the third-party cap, upgrade them. Standard seat overages are costly. Premium is more cost-effective. The pools are siloed. Unused first-party tokens on Standard cannot offset third-party overages. This rigidity demands precise role mapping.
Monitor the usage dashboard weekly. Do this during the first month. Identify developers draining third-party allowance by week two. These are your Premium candidates. Conversely, some developers rarely touch external APIs. They can remain on Standard seats. Productivity is not impacted. This targeted approach maximizes developer productivity. It controls AI coding costs.
How can teams optimize costs under the new dual-pool system?
We can reduce expenses significantly. Prioritize first-party models for routine tasks. This preserves the limited third-party API allowance. Use it for complex reasoning. Save it for specialized external model requirements.
Prioritizing native models for routine tasks
The split pricing rewards alignment with native capabilities. The first-party pool includes Auto, Composer 2.5, and Grok 4.5. Direct most daily editing to these tools https://forum.cursor.com/t/teams-first-party-model-pool/165385. These models are optimized for the IDE. They do not drain the third-party budget. Make Cursor Auto the default for boilerplate. Use it for simple logic fixes. Keep the expensive external API pool available. Save it for high-stakes architectural decisions.
This requires a shift in habits. Train your team to recognize needs. Does a task truly demand an external model? Claude or GPT might be overkill. Native models provide sufficient quality for many operations. The effective cost is lower for the organization. Cursor states this lowers costs for 90% of teams. Most usage patterns can adapt https://dreaming.press/posts/cursor-teams-two-usage-pools-premium-seat.html. Encourage developers to start with Composer 2.5. Switch providers only if initial output fails. This disciplined routing prevents accidental exhaustion. Do not burn through the third-party allowance early.
Context management strategies for expensive API calls
Precise context management is critical for cost control. We must be careful with third-party models. Third-party API model usage is charged at public list prices. The Cursor Token Rate is added https://cursor.com/docs/account/teams/pricing. Every token counts. Avoid dumping entire codebases into the context window. Manually select only relevant files. Include only functions needed for the prompt. This selective inclusion reduces token burn. It extends the life of our third-party pool.
Establish internal guidelines for context hygiene. Developers should close unused tabs. Clear chat history before starting complex queries. Use explicit file references in prompts. This helps the AI focus. It avoids pulling in irrelevant dependencies. Monitor the usage dashboard regularly. Track consumption rates for both pools. Intervene if you see a spike in third-party usage. Provide targeted training on prompt efficiency. Proactive management keeps us within limits. It avoids unexpected overage charges. The pools are siloed. We cannot borrow from one to cover the other. Disciplined usage is our only safeguard. The vexp tool can help visualize these consumption patterns across your team.
Frequently Asked Questions
Does Cursor Auto use the first-party or third-party pool?
What happens when I exceed my third-party API allowance?
Can I upgrade just the third-party pool without changing seats?
Is Composer 2.5 considered a first-party model?
How do I check my current usage for both pools?
What defines the first-party models pool?
What is included in the third-party API allowance?
How does the dual-pool system affect pricing?
Nicola
Developer and creator of vexp — a context engine for AI coding agents. I build tools that make AI coding assistants faster, cheaper, and actually useful on real codebases.
Keep reading
Related articles

AI Code Maintainability Decline 2026: Data, Causes, and Fixes
Discover 2026 data on AI code maintainability decline, including AI technical debt, write-only code, and code churn metrics. Learn fixes to prevent software quality

Uber Caps AI Spend After Burning 2026 Budget on Claude Code
Uber burned its 2026 AI budget in four months on Claude Code, enforcing a $1,500 monthly cap per employee. Learn token optimization strategies to avoid overspend.

MCP 2026-07-28 Spec: Stateless Core & Migration Guide
Learn about the MCP 2026-07-28 spec with a stateless core, breaking changes, and a migration guide. Optimize token usage and scale AI apps easily.