Assistant Best Practices
This guide compiles proven patterns, strategies, and lessons learned from building successful assistants. Follow these practices to create assistants that perform reliably, cost-effectively, and delight users.Design Principles
Start Simple, Add Complexity
The Progressive Enhancement Approach
- Single, clear purpose
- 1-2 essential tools
- Basic instructions
- Simple happy path
- Test with real users
- Add edge case handling
- Refine instructions based on feedback
- Optimize tool usage
- Add advanced features
- More tools as needed
- Sophisticated error handling
- Performance optimization
- Faster to get live
- Easier debugging
- Clear performance baseline
- Incremental improvement
Single Responsibility Principle
Each assistant should have one clear purpose.- Good: Focused Assistants
- Bad: Swiss Army Knife Assistant
- ✅ Clear purpose
- ✅ Easier to optimize
- ✅ Simpler instructions
- ✅ Better performance
- ✅ Easier to debug
- Assistant instructions exceed 2,000 words
- Assistant has 10+ tools
- Performance is inconsistent
- Different user groups with different needs
- Clear logical separation of concerns
Instruction Writing
Be Obsessively Specific
Vague instructions produce inconsistent results. Specificity drives performance.Define Success Clearly
Define Success Clearly
Quantify Everything
Quantify Everything
Show, Don't Tell
Show, Don't Tell
Spell Out Edge Cases
Spell Out Edge Cases
Tool Management
Tool Selection Strategy
The 80/20 Rule for Tools
- Primary data source (knowledge base, CRM, database)
- Most common action tool (create ticket, process refund)
- Secondary data sources
- Additional action tools
- Advanced features
- Nice-to-have integrations
Tool Usage Patterns
Search Before You Answer
Search Before You Answer
Verify Before You Act
Verify Before You Act
Enrich Before You Qualify
Enrich Before You Qualify
Escalate When Uncertain
Escalate When Uncertain
Performance Optimization
Token Efficiency
Right-Size Context Windows
Right-Size Context Windows
- Check actual token usage in logs
- Are you consistently near the limit? → Increase
- Are you using < 50% of limit? → Decrease
- Enable Smart Context (reduces tokens automatically)
- Limit message history to what’s actually needed
- Trim verbose tool descriptions
- Use concise instructions
- Simple Q&A: 16K tokens
- Standard assistants: 50K tokens
- Complex assistants: 100K tokens
- Document processing: 128K+ tokens
Optimize Tool Descriptions
Optimize Tool Descriptions
Reduce Unnecessary Tool Calls
Reduce Unnecessary Tool Calls
- Tell the assistant clearly when a tool call is (and isn’t) needed
- Avoid instructions that encourage re-checking or re-verifying the same thing repeatedly
- Review conversations for repeated or redundant tool calls and tighten instructions when you see them
Cache Common Queries
Cache Common Queries
- “What are your hours?” (asked 100x/day)
- “What’s your return policy?” (asked 50x/day)
- Common product questions
- Identify top 20 repeated questions
- Pre-generate high-quality responses
- Store in fast-access cache
- Return cached response when matched
- Fall back to assistant for unique queries
- Instant responses (< 100ms)
- Zero token cost for cached hits
- Consistent quality
- Reduced API load
Cost Management
Choose the Right Model
Choose the Right Model
- Simple classification or routing: pick the fastest, cheapest model in the picker.
- Standard automation and everyday conversation: a mid-tier model is usually enough.
- Complex reasoning or high-stakes tasks: step up to a larger, more capable model only when quality actually demands it.
- Task: Simple lead qualification (company size, industry match)
- Start with: the smallest, cheapest model in the picker
- Upgrade only if: task complexity requires deeper reasoning
- Performance: a small model handles most qualification tasks well
Watch Your Spend
Watch Your Spend
- Cost per assistant run
- Cost per day/week/month
- Token usage per assistant
- Most expensive assistants
- Unusual spikes
- Which assistants cost the most?
- Can any be optimized?
- Are costs justified by value?
Optimize Retries
Optimize Retries
- Better input validation
- Clearer instructions
- An output schema to guide the model toward a consistent response format
- Better error handling
- Testing edge cases
Quality Assurance
Testing Checklist
Before rolling out to production, test:Happy Path Scenarios
- 10 typical, straightforward interactions
- Verify assistant responds correctly
- Check tool usage is appropriate
- Confirm output format
Edge Cases
- 5-10 unusual but possible scenarios
- Past-policy refund requests
- Missing data
- Tool failures
- Ambiguous requests
Error Conditions
- Invalid inputs
- Tool timeouts
- Authentication failures
- Rate limit errors
- Malformed data
Adversarial Cases
- Attempts to break role
- Extremely long inputs
- Nonsense queries
- Rapid-fire questions
- Contradictory requests
Performance
- Response time acceptable?
- Token usage reasonable?
- Cost per interaction acceptable?
- No memory leaks or hangs?
User Experience
- Tone is appropriate?
- Responses are helpful?
- Escalation works smoothly?
- Overall experience positive?
Monitoring in Production
Track Success Metrics
Track Success Metrics
- % inquiries resolved without escalation
- Average response time
- Customer satisfaction score
- Tool usage accuracy
- Cost per resolution
- % leads qualified automatically
- Qualification accuracy (validated by sales)
- Meeting booking rate
- Time saved per lead
- Cost per qualified lead
- Week over week improvement?
- Seasonal variations?
- Degradation after changes?
Review Conversations Weekly
Review Conversations Weekly
- 10 random conversations
- 5 escalated conversations
- 5 low-satisfaction conversations
- 5 high-satisfaction conversations
- Instruction following
- Tool usage appropriateness
- Tone and communication quality
- Edge cases not yet handled
- Opportunities for improvement
Security Best Practices
Protect Customer Data
Protect Customer Data
Prevent Prompt Injection
Prevent Prompt Injection
Secure Tool Access
Secure Tool Access
- Basic authentication sufficient
- Minimal risk
- Require strong authentication
- Implement monetary/scope limits
- Add human approval for high-value actions
- Log all actions
- Set up alerts for unusual activity
Audit Logs
Audit Logs
- Timestamp
- User identifier (hashed/anonymized if needed)
- Assistant used
- Input prompt
- Assistant response
- Tools called (with parameters)
- Errors encountered
- Token usage
- Cost
- Security audits
- Debugging issues
- Performance analysis
- Compliance reporting
- Fraud detection
Common Pitfalls to Avoid
Over-Engineering Before Validation
Over-Engineering Before Validation
- ❌ Build assistant with 15 tools and 5,000-word instructions on day 1
- ✅ Build assistant with 2 tools and 500-word instructions. Test. Iterate.
Ignoring Real User Feedback
Ignoring Real User Feedback
- ❌ “I think users want X” → Build X
- ✅ Review 50 conversations → Users actually need Y → Build Y
Not Handling Tool Failures
Not Handling Tool Failures
Vague Success Criteria
Vague Success Criteria
- ❌ “Assistant should help customers”
- ✅ “Assistant should: 1) Resolve 75% of inquiries without escalation, 2) Response time < 30 seconds, 3) CSAT > 4.5/5”
Not Versioning Instructions
Not Versioning Instructions
- Keep instructions in version control (Git)
- Document changes in commits
- Tag major versions
- Can A/B test versions
- Can roll back if new version performs worse
Optimizing Prematurely
Optimizing Prematurely
- Make it work: Basic functionality, correct behavior
- Make it good: Refine quality, handle edge cases
- Make it efficient: Optimize tokens, cost, speed
Rollout Strategy
Phased Rollout
Internal Testing (Week 1)
- Use it with your own team only
- Test with real scenarios
- Gather feedback from colleagues
- Fix critical issues
Limited Users (Week 2-3)
- Share it with a small group of real users
- Monitor closely
- Rapid iteration based on feedback
- Validate success metrics
Gradual Rollout (Week 4-6)
- Add more users as confidence grows
- Watch for degradation or issues
- Compare metrics against your baseline
- Adjust as needed
Full Availability (Week 7+)
- Make it available to everyone who needs it
- Continue monitoring
- Iterate based on data
- Celebrate success! 🎉
Continuous Improvement
Weekly Optimization Routine
Monday: Review Metrics
- Check success metrics vs. targets
- Identify trends (improving or degrading?)
- Flag anomalies
Tuesday: Review Conversations
- Sample 10-20 conversations
- Look for improvement opportunities
- Note edge cases not handled well
Wednesday: Identify Improvements
- Based on metrics and conversations
- Prioritize by impact and effort
- Select 1-2 improvements to implement
Thursday: Implement & Test
- Update instructions
- Test changes thoroughly
- Prepare A/B test if significant change
Friday: Save & Monitor
- Save improvements
- Watch metrics closely
- Gather early feedback
Monthly Deep Dive
Once per month, conduct a thorough review:Performance Analysis
Performance Analysis
- Review all metrics for the month
- Compare to previous months
- Identify trends
- Calculate ROI
Cost Analysis
Cost Analysis
- Total spend for the month
- Cost per interaction
- Most expensive assistants
- Optimization opportunities
- ROI calculation
User Satisfaction
User Satisfaction
- CSAT trends
- Qualitative feedback themes
- Feature requests
- Pain points
- Success stories
Technical Health
Technical Health
- Error rates
- Tool reliability
- Response times
- Token usage
- Areas for technical improvement
Strategic Planning
Strategic Planning
- What’s working well?
- What needs improvement?
- New use cases to explore?
- Tools to add or remove?
- Next quarter priorities
Success Stories & Patterns
What Great Assistants Have in Common
Analyzing top-performing assistants reveals common patterns:Crystal Clear Purpose
Specific Instructions
Rich Examples
Edge Case Coverage
Right-Sized Tooling
Clear Escalation Path
Continuous Iteration
Measurable Success
Quick Reference Checklist
Use this checklist when building or optimizing assistants:Design
- Assistant has single, clear purpose
- Instructions are specific and actionable
- 2-3 complete example scenarios included
- Top 5-10 edge cases handled explicitly
- Success metrics defined clearly
Tools
- Only essential tools connected
- Each tool has clear usage guidelines
- Tool authentication tested and working
- Escalation path defined for tool failures
Configuration
- Appropriate model selected for task complexity
- Smart Context enabled
- Message History Limit sized to how much conversation history the assistant actually needs
Testing
- 10+ happy path scenarios tested
- 5+ edge cases tested
- Error conditions tested
- Performance acceptable (speed and cost)
- User experience validated
Rollout
- Phased rollout plan in place
- Previous instructions saved somewhere you can restore from
- Spend and usage reviewed regularly
Maintenance
- Weekly review scheduled
- Monthly deep dive planned
- Feedback collection process in place
- Continuous improvement mindset