# Advanced AI APIs, Features, & Integration Tools :u-color-mode-image{alt="APIpie" dark="/img/docs/apipie-logo-full.png" light="/img/docs/apipie-logo-full-light.png"} \:: Unlock the power of AI with APIpie.ai, your AI super aggregator. Our unified API simplifies your access to an extensive array of AI services from leading providers, enabling both cost-effective and latency-optimized solutions. With one subscription, leverage advancements in language, vision, embeddings, and more - designed for seamless integration and optimized for both developers and businesses. Start building with APIpie.ai and transform your applications with cutting-edge AI technology. # AI Comparison Guide: Overview of LLM Providers ![Models Overview Banner](https://apipie.ai/img/docs/models/Overview.svg){width="100%"} # AI Models Overview Welcome to APIpie.ai's comprehensive AI model catalog. We provide access to hundreds of models across multiple providers, all through a unified API interface. This overview helps you understand our model offerings and choose the right models for your needs. ## Model Comparison Matrix | Model Family | Best For | Context Window | Key Features | Specialized Capabilities | | ------------ | ------------------------------ | -------------- | ------------------------------------ | -------------------------------------- | | GPT-4 | Enterprise, Complex Tasks | 32K-128K | Tool use, Vision, Advanced reasoning | Function calling, JSON mode | | Claude 3 | Long-form, Analysis | 200K+ | Long context, Detailed analysis | Multi-turn reasoning, Research | | Gemini | Multimodal, Code | 32K+ | Vision, Code generation | Real-time processing, Math | | Llama 3 | Open source, Custom deployment | 8K-128K | Customizable, Local deployment | Domain adaptation | | Mistral | Efficient, Specialized | 8K-32K | Task-specific models | Mixture of experts | | Cohere | Enterprise, Embeddings | 32K+ | Multilingual, Custom training | Document processing | | AI21 | Domain-specific tasks | 8K-32K | Specialized models | Academic writing | | Amazon Titan | Enterprise security | 8K-32K | Compliance features | Regulated industries | | Qwen | Multilingual, Asian languages | 8K-32K | Chinese optimization | Cross-lingual tasks | | Perplexity | Research, RAG | 32K+ | Citation support | Knowledge integration, Internet Search | ## Performance Benchmarks ### Language Understanding | Model | MMLU Score | HumanEval | GSM8K | | -------------- | ---------- | --------- | ----- | | GPT-4 | 86.4% | 67% | 92% | | Claude 3 | 85.2% | 71% | 88% | | Gemini Pro | 83.7% | 63% | 85% | | Mistral Large | 82.6% | 59% | 81% | | Cohere Command | 81.5% | 58% | 79% | | AI21 Jurassic | 80.8% | 56% | 77% | | Titan Express | 79.9% | 54% | 75% | | Qwen Max | 81.2% | 57% | 78% | | Perplexity | 82.1% | 60% | 80% | ### Image Generation | Model | Resolution | Style Control | Speed | | ---------------- | ---------- | ------------- | ------ | | DALL·E 3 | 1024x1024 | High | Fast | | Stable Diffusion | 1024x1024 | Very High | Medium | | Midjourney | 1024x1024 | High | Fast | ## Major AI Providers ### [OpenAI](https://apipie.ai/docs/models/openai) - **GPT-4 Series** - Latest GPT-4 models including Turbo and Vision variants - **GPT-3.5 Series** - Cost-effective GPT-3.5 models - **Text Embeddings** - Ada-002 and Text Embedding v3 models - **Image Generation** - DALL·E 2 and DALL·E 3 - **Audio Models** - Whisper and TTS models ### [Anthropic](https://apipie.ai/docs/models/claude) - **Claude Series** - Including Claude 3 Opus, Sonnet, and Haiku - **Claude Instant** - Fast, cost-effective variant - Detailed documentation on capabilities and best practices ### [Google](https://apipie.ai/docs/models/google) - **Gemini Series** - Pro, Ultra, and Vision models - **PaLM Models** - Chat and code-specialized variants - **Embedding Models** - Text and multilingual embeddings ### [Meta](https://apipie.ai/docs/models/llama) - **Llama Series** - Llama 2 and Llama 3 models - **Vision Models** - Multimodal capabilities - **LlamaGuard** - Safety-focused models ### [Mistral](https://apipie.ai/docs/models/mistral) - **Mistral Models** - From small to large variants - **Mixtral** - Mixture of experts architecture - Advanced instruction-following capabilities ### [AI21 Labs](https://apipie.ai/docs/models/ai21) - **Jurassic Models** - Various sizes and specializations - **Specialized Models** - Task-specific variants - Multilingual capabilities ### [Cohere](https://apipie.ai/docs/models/cohere) - **Command Series** - Different sizes and capabilities - **Embedding Models** - Multilingual and specialized - Enterprise-grade reliability ### [Amazon](https://apipie.ai/docs/models/amazon) - **Titan Models** - Text and image generation - **Embedding Models** - Text and multimodal - Enterprise security features ## Specialized Models ### Image Generation - **Stable Diffusion** - Various fine-tuned versions - **Midjourney-style** - Artistic and creative models - **Realistic** - Photorealistic generation models ### Code Generation - **CodeLlama** - Specialized code models - **StarCoder** - Multiple programming languages - **Specialized variants** - Language-specific models ### Embedding Models - **Sentence Transformers** - Various sizes and specializations - **BGE Models** - Multilingual capabilities - **Domain-specific** - Task-optimized embeddings ### Voice Synthesis - **ElevenLabs** - High-quality voice generation - **Multilingual** - Support for multiple languages - **Custom Voices** - Voice cloning capabilities ## Model Selection Guide ### Factors to Consider 1. **Task Requirements** - Language understanding - Code generation - Image creation - Voice synthesis 2. **Performance Needs** - Response time - Accuracy - Cost efficiency - Scalability 3. **Technical Constraints** - API quotas - Token limits - Integration complexity - Security requirements 4. **Cost Considerations** - Per-token pricing - Volume discounts - Additional features ## Integration Support ### API Compatibility - OpenAI-compatible endpoints - Standardized request formats - Consistent response structures - Easy migration paths ### Features - Model pooling for reliability - Automatic failover - Cost optimization - Usage monitoring ### Security - Enterprise-grade security - Data privacy controls - Compliance features - Audit capabilities ## Frequently Asked Questions About AI Models ### Which AI model is best for enterprise use? The best model depends on your specific needs. GPT-4 and Claude 3 excel at general tasks, while specialized models like Cohere and AI21 offer unique capabilities for specific use cases. ### How do I choose between different AI providers? Consider factors like: - Performance requirements - Cost constraints - Security needs - Integration complexity - Specific features needed ### What's the difference between GPT-4 and Claude 3? While both are powerful models: - GPT-4 offers broader integration options and tool use - Claude 3 excels at longer context and detailed analysis - Both provide strong security and reliability ### How do image generation models compare? - DALL·E 3: Best for realistic and creative images - Stable Diffusion: More customizable and cost-effective - Midjourney-style: Artistic and stylized outputs ### What are the cost considerations for AI models? Costs vary by: - Per-token pricing - Volume discounts - Additional features - Integration complexity - Infrastructure requirements ### How can I ensure reliable AI model performance? - Use model pooling for redundancy - Implement proper error handling - Monitor usage and performance - Choose appropriate model sizes - Consider fallback options ### What security features should I look for? - Data privacy controls - Compliance certifications - Audit capabilities - Access controls - Data retention policies ### How do embedding models differ? - Size and performance trade-offs - Multilingual capabilities - Domain specialization - Integration complexity - Cost considerations ## Getting Started 1. Browse our detailed model documentation for specific capabilities 2. Explore our [features](https://apipie.ai/docs/features) 3. Read the [API Reference](https://apipie.ai/docs/api/introduction){rel=""nofollow""} 4. Monitor usage and performance in the [dashboard](https://apipie.ai/dashboard){rel=""nofollow""} Ready to get started? [Sign up for APIpie.ai](https://apipie.ai/profile/auth/register){rel=""nofollow""} to access our comprehensive suite of AI models and features. Compare models, test capabilities, and find the perfect solution for your needs. # Amazon AI Guide: Titan & Nova Series Comparison ![Amazon Bedrock](https://apipie.ai/img/docs/models/AWS.png){width="50%"} ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Description** [Amazon's Titan and Nova models](https://aws.amazon.com/ai/generative-ai/nova/){rel=""nofollow""} represent AWS's proprietary family of foundation models, available through [Amazon Bedrock](https://aws.amazon.com/bedrock/){rel=""nofollow""} - AWS's unified AI model platform. These models are designed to handle a variety of tasks including text generation, image generation, and embeddings, offering enterprise-grade reliability and performance through [APIpie's routing system](https://apipie.ai/docs/features/routing). ### **Amazon Bedrock as a Provider** While this guide focuses on Amazon's own Titan and Nova models, Amazon Bedrock serves as a comprehensive AI model platform hosting both Amazon's proprietary models and those from other leading providers. Through APIpie's unified API, you can seamlessly access: - Amazon's proprietary models (Titan, Nova) - Third-party provider models hosted on Bedrock - Comprehensive monitoring and cost tracking - Automatic model routing and fallback APIpie simplifies Bedrock integration by providing: - Single API endpoint for all Bedrock models - Unified pricing and performance monitoring - Automated provider selection and failover - Consistent response format across all models For a complete list of available models and providers, see our [Models Route](https://apipie.ai/docs/features/models). #### **Model Families** **Titan Models:** - **Text Series:** Ranging from lite to premier versions, optimized for different performance and resource needs - **Embedding Models:** Specialized for text and image embeddings, ideal for search and recommendation systems - **Image Generation:** Advanced models for creating and manipulating images **Nova Models:** - **Text Processing:** Specialized in handling extremely long contexts up to 300K tokens - **Performance Tiers:** From micro to pro versions, balancing efficiency and capability - **Specialized Variants:** Canvas and Reel versions for specific use cases #### **Key Features** - **Extended Token Capacity:** Models support context lengths from 4K to 300K tokens for various processing needs. - **Diverse Capabilities:** Text generation, image generation, and embedding models available. - **Enterprise Focus:** Built for production-grade applications with AWS's reliability. - **Optimized Performance:** Models ranging from lightweight to high-performance versions. --- ### **Model Comparison and Monitoring** When choosing between Titan and Nova models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (Bedrock, OpenRouter) - Latency metrics and historical performance data - Response time comparisons between different model versions **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons and pricing trends across different providers **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics This monitoring system helps users: - Compare costs and pricing across different Titan and Nova models - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all Amazon models. :: ### **Model Feature Comparison** When choosing between Titan and Nova models, consider these key differences: **Titan Models:** - Best for: General-purpose text generation, image tasks, and embeddings - Strengths: - More versatile with text, image, and embedding capabilities - Lower latency for standard tasks - Optimized for production workloads - Use when: - You need image generation or embedding capabilities - Working with standard context lengths - Requiring fast response times **Nova Models:** - Best for: Long-form content and complex document processing - Strengths: - Much larger context windows (up to 300K tokens) - Specialized for long-form content - Advanced text processing capabilities - Use when: - Processing very long documents - Needing extensive context for responses - Working with complex, multi-part texts --- ### **Model List** #### **Amazon's Titan & Nova Models** While Amazon Bedrock hosts various provider models, this section focuses on Amazon's own Titan and Nova series. For a complete list of all available models through Bedrock and other providers, see our [Models Route](https://apipie.ai/docs/features/models). ::note For detailed information about model capabilities and performance, visit the [Amazon Titan Documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/titan.html){rel=""nofollow""}. :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Providers** | **Type** | | ---------------------------------- | -------------- | ------------------- | ------------------- | --------- | | titan-tg1-large | 32,000 | 32,000 | Bedrock | LLM | | titan-text-lite-v1 | 4,000 | 4,000 | Bedrock | LLM | | titan-text-express-v1 | 8,000 | 8,000 | Bedrock | LLM | | titan-text-premier-v1 | 32,000 | 32,000 | Bedrock | LLM | | olympus-premier-v1 | - | - | Bedrock | LLM | | titan-embed-g1-text-02 | 8,192 | 8,192 | Bedrock | Embedding | | titan-embed-text-v1 | 8,192 | 8,192 | Bedrock | Embedding | | titan-embed-text-v2 | - | - | Bedrock | Embedding | | titan-image-generator-v1 | - | - | Bedrock | Image | | titan-image-generator-v2 | - | - | Bedrock | Image | | titan-embed-image-v1 | - | - | Bedrock | Image | | titan-image-generator-v1\_premium | - | - | EdenAI | Image | | titan-image-generator-v1\_standard | - | - | EdenAI | Image | | nova-pro-v1 | 300,000 | 5,120 | Bedrock, OpenRouter | LLM | | nova-lite-v1 | 300,000 | 5,120 | Bedrock, OpenRouter | LLM | | nova-micro-v1 | 128,000 | 5,120 | Bedrock, OpenRouter | LLM | | nova-canvas-v1 | - | - | Bedrock | LLM | | nova-reel-v1 | - | - | Bedrock | LLM | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a Titan model. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "bedrock", "model": "titan-text-premier-v1", "max_tokens": 150, "messages": [ { "role": "user", "content": "Can you explain how photosynthesis works?" } ] }' ``` --- ### Response Example The expected response structure for a Titan model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "bedrock", "model": "titan-text-premier-v1", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Photosynthesis is the process by which plants convert sunlight into chemical energy. Here's how it works:\n\n1. **Light Absorption**: Plants capture sunlight using chlorophyll in their leaves\n\n2. **Water and CO2**: They take in water through roots and carbon dioxide through leaf pores\n\n3. **Chemical Reaction**: Using sunlight's energy, they convert H2O and CO2 into glucose and oxygen:\n 6CO2 + 6H2O + light → C6H12O6 + 6O2\n\nThis process produces food for the plant and releases oxygen as a byproduct." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.00225, "latency_ms": 3100 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify "bedrock" as the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model Selection:** - Use Titan models for general text generation and embeddings - Use Nova models for extended context processing - See [Models Guide](https://apipie.ai/docs/features/models#fetching-models) for the complete list - **Max Tokens:** Set according to model capacity (varies by model) - **Messages:** Format your request following the [message formatting guide](https://apipie.ai/docs/features/completions) --- ### **Applications and Integrations** **Titan Applications:** - **Text Generation:** - Content creation and chat applications - Premier variant (32K context) for complex tasks - Express variant (8K context) for balanced performance - Lite variant (4K context) for efficient processing - **Image Generation:** - Creating and editing images - Visual content generation - Image manipulation and enhancement - **Embeddings:** - Semantic search implementation - Document similarity analysis - Content recommendation systems - Cross-modal applications with image embeddings **Nova Applications:** - **Extended Context Processing:** - Document analysis up to 300K tokens - Long-form content generation - Complex document summarization - **Specialized Processing:** - Canvas variant for visual-heavy content - Reel variant for sequential content - Micro variant for efficient processing of medium-length content (128K tokens) **Integration Options:** - Try the models with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI) --- ### **Ethical Considerations** Amazon's models are designed with responsible AI principles. Users should implement appropriate safeguards and consider potential biases. For guidance, see AWS's [Responsible AI Guidelines](https://aws.amazon.com/machine-learning/responsible-ai/){rel=""nofollow""}. --- ### **Licensing** Titan and Nova models are available through AWS Bedrock. Usage is subject to [AWS Service Terms](https://aws.amazon.com/service-terms/){rel=""nofollow""}. For detailed information, consult the [Amazon Bedrock documentation](https://docs.aws.amazon.com/bedrock/){rel=""nofollow""}. ::tip Try out the Titan and Nova models in APIpie's various \[supported integrations.]\(/docs/Sandbox/Chat) :: # Claude AI by APIpie: Overview & Key Features ![Anthropic Claude](https://apipie.ai/img/docs/models/Claude.png){width="100%"} ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Description** The [Claude Series](https://www.anthropic.com/claude){rel=""nofollow""} represents Anthropic's family of state-of-the-art large language models. These models, developed using [Constitutional AI](https://www.anthropic.com/research){rel=""nofollow""}, leverage cutting-edge technology to deliver exceptional performance in natural language processing, instruction-following tasks, and extended-context interactions. The models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). #### **Key Features** - **Extended Token Capacity:** Models support context lengths from 8K to 200K tokens for handling various text processing needs. - **Multi-Provider Availability:** Accessible across platforms like [OpenRouter](https://openrouter.ai/){rel=""nofollow""}, [EdenAI](https://www.edenai.co/){rel=""nofollow""}, [Anthropic](https://www.anthropic.com/){rel=""nofollow""}, and [Amazon Bedrock](https://aws.amazon.com/bedrock/){rel=""nofollow""}. - **Diverse Applications:** Optimized for chat, instruction-following, analysis, writing, and multimodal tasks. - **Constitutional AI:** Built with advanced safety measures and ethical considerations. - **Multimodal Capabilities:** Claude-3 models can understand and analyze images alongside text. --- ### **Model Comparison and Monitoring** When choosing between Claude models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (Anthropic, OpenRouter, Bedrock, EdenAI) - Latency metrics and historical performance data - Response time comparisons between different model versions **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons and pricing trends across different providers - Usage-based cost optimization recommendations **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics This monitoring system helps users: - Compare costs and pricing across different Claude models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all Claude models. :: ### **Model List in the Claude Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::info For information about model performance and benchmarks, see the \[Claude Technical Report]\(https\://www-files.anthropic.com/production/images/Model-Card-Claude-3.pdf). The Claude-3 and Claude-3.5 series (Opus, Sonnet, Haiku) represent the latest generations with enhanced capabilities including multimodal understanding and increased response token limits up to 8,192 tokens, while the Claude-2 series offers robust performance for text-based tasks. :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Provider** | **Type** | | ---------------------------- | -------------- | ------------------- | ------------ | -------- | | claude-3-opus | 200,000 | 4,096 | Anthropic | LLM | | claude-3-5-sonnet | 200,000 | 4,096 | Anthropic | LLM | | claude-3-opus | 200,000 | 4,096 | OpenRouter | LLM | | claude-3-sonnet | 200,000 | 4,096 | OpenRouter | Vision | | claude-3-haiku | 200,000 | 4,096 | OpenRouter | Vision | | claude-3-5-sonnet | 200,000 | 8,192 | OpenRouter | LLM | | claude-3-5-haiku | 200,000 | 8,192 | OpenRouter | LLM | | claude-3-5-haiku-20241022 | 200,000 | 8,192 | OpenRouter | LLM | | claude-2.0 | 100,000 | 4,096 | OpenRouter | LLM | | claude-2.1 | 200,000 | 4,096 | OpenRouter | LLM | | claude-3-sonnet | 200,000 | 4,096 | Bedrock | Vision | | claude-3-haiku | 200,000 | 4,096 | Bedrock | Vision | | claude-3-opus-20240229-v1 | 200,000 | 4,096 | Bedrock | LLM | | claude-3-5-sonnet | 200,000 | 4,096 | Bedrock | LLM | | claude-3-5-haiku-20241022-v1 | 200,000 | 4,096 | Bedrock | LLM | | claude-2 | 200,000 | 4,096 | Bedrock | LLM | | claude-instant-1 | 100,000 | 4,096 | Bedrock | LLM | | claude-3-haiku | 200,000 | 4,096 | EdenAI | Vision | | claude-3-sonnet | 200,000 | 4,096 | EdenAI | Vision | | claude-3-5-sonnet | 200,000 | 4,096 | EdenAI | LLM | | claude-2 | 200,000 | 4,096 | EdenAI | LLM | | claude-instant-1 | 100,000 | 4,096 | EdenAI | LLM | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a model from the **Claude Series**, such as `claude-3-opus`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "anthropic", "model": "claude-3-opus", "max_tokens": 150, "messages": [ { "role": "user", "content": "Can you explain how photosynthesis works?" } ] }' ``` --- ### Response Example The expected response structure for the Claude model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "anthropic", "model": "claude-3-opus", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Photosynthesis is the process by which plants convert sunlight into chemical energy. Here's how it works:\n\n1. **Light Absorption**: Plants capture sunlight using chlorophyll in their leaves\n\n2. **Water and CO2**: They take in water through roots and carbon dioxide through leaf pores\n\n3. **Chemical Reaction**: Using sunlight's energy, they convert H2O and CO2 into glucose and oxygen:\n 6CO2 + 6H2O + light → C6H12O6 + 6O2\n\nThis process produces food for the plant and releases oxygen as a byproduct." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.00225, "latency_ms": 3100 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the Claude Series suited to your task. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/features/completions). This example demonstrates how to seamlessly query models from the Claude Series for conversational or instructional tasks. --- ### **Applications and Integrations** - **Conversational AI:** Powering chatbots, virtual assistants, and other dialogue-based systems. Try it with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI). - **Analysis and Research:** Leveraging Claude's strong analytical capabilities for data analysis and research tasks. - **Content Creation:** Using Claude's writing abilities for content generation and editing. - **Extended Context Tasks:** Processing long documents with models supporting up to 200K tokens. Learn more in our [Models Guide](https://apipie.ai/docs/features/models). - **Educational Support:** Providing detailed explanations and tutoring across various subjects. - **Multimodal Analysis:** Using Claude-3's ability to understand and analyze images alongside text for richer interactions. - **Vision Tasks:** Leveraging image understanding capabilities for tasks like visual analysis, chart interpretation, and document processing. --- ### **Ethical Considerations** Claude models are built with Constitutional AI principles to ensure safe and ethical behavior. Users should still implement appropriate safeguards and consider potential biases in model outputs. For guidance on responsible AI usage, see Anthropic's [Safety Practices](https://support.anthropic.com/en/articles/8106465-our-approach-to-user-safety){rel=""nofollow""}. --- ### **Licensing** The Claude Series is available through Anthropic's API and partner platforms. Usage is subject to Anthropic's [Terms of Service](https://www.anthropic.com/legal/terms){rel=""nofollow""} and respective provider agreements. For detailed licensing information, consult [Anthropic's documentation](https://docs.anthropic.com/){rel=""nofollow""} and respective hosting providers. ::tip Try out the Claude models in APIpie's various \[supported integrations.]\(/docs/Sandbox/Chat) :: # Discover Cohere Models: API for AI Applications ![Cohere](https://apipie.ai/img/docs/models/Cohere.png){width="100%"} ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Description** [Cohere](https://cohere.com/){rel=""nofollow""} is a leading provider of enterprise-grade language AI models. Their [Command series](https://docs.cohere.com/docs/command){rel=""nofollow""} represents state-of-the-art models optimized for production environments, offering exceptional performance in natural language understanding, generation, and task completion. These models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). #### **Key Features** - **Extended Token Capacity:** Models support context lengths from 4K to 128K tokens, with [Command-R models](https://docs.cohere.com/docs/command-r){rel=""nofollow""} optimized for processing long documents. - **Multi-Provider Availability:** Accessible through [Amazon Bedrock](https://aws.amazon.com/bedrock){rel=""nofollow""}, [OpenRouter](https://openrouter.ai/){rel=""nofollow""}, [EdenAI](https://www.edenai.co){rel=""nofollow""}. - **Diverse Applications:** Optimized for chat, [instruction-following](https://docs.cohere.com/docs/prompt-engineering){rel=""nofollow""}, and [text generation](https://docs.cohere.com/docs/text-generation){rel=""nofollow""} tasks. - **Enterprise Focus:** Models designed for production-ready deployment with consistent performance and [enterprise-grade security](https://cohere.com/security){rel=""nofollow""}. - **Advanced Capabilities:** Features include [semantic search](https://docs.cohere.com/docs/semantic-search){rel=""nofollow""}, text classification, and content moderation. --- ### **Model Comparison and Monitoring** When choosing between Cohere models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (Bedrock, OpenRouter, EdenAI) - Latency metrics and historical performance data - Response time comparisons between different model versions **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons and pricing trends across different providers - Usage-based cost optimization recommendations **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics This monitoring system helps users: - Compare costs and pricing across different Cohere models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all Cohere models. :: ### **Model List in the Cohere Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::info For detailed performance metrics and benchmarks, see the \[Cohere Model Performance]\(https\://artificialanalysis.ai/providers/cohere) documentation and \[Command Model Cards]\(https\://docs.cohere.com/v2/docs/models). :: #### Language Models (LLMs) | **Model Name** | **Max Tokens** | **Response Tokens** | **Providers** | **Subtype** | | ---------------------- | -------------- | ------------------- | ------------------ | ----------- | | command | 4,096 | 4,000 | OpenRouter, EdenAI | Chat | | command-text-v14 | 4,000 | 4,000 | Bedrock | Chat | | command-light-text-v14 | 4,000 | 4,000 | Bedrock | Chat | | command-r-v1 | 128,000 | 4,000 | Bedrock | Chat | | command-r-plus-v1 | 128,000 | 4,000 | Bedrock | Chat | | command-r | 128,000 | 4,000 | OpenRouter, EdenAI | Chat | | command-r-plus | 128,000 | 4,000 | OpenRouter, EdenAI | Chat | | command-r-08-2024 | 128,000 | 4,000 | OpenRouter | Chat | | command-r-plus-08-2024 | 128,000 | 4,000 | OpenRouter | Chat | | command-r-plus-04-2024 | 128,000 | 4,000 | OpenRouter | Chat | | command-r-03-2024 | 128,000 | 4,000 | OpenRouter | Chat | | command-r7b-12-2024 | 128,000 | 4,000 | OpenRouter | Chat | | command-light | 4,096 | 4,000 | EdenAI | Chat | | command-light-nightly | 4,096 | 4,000 | EdenAI | Chat | | command-nightly | 4,096 | 4,000 | EdenAI | Chat | #### Embedding Models | **Model Name** | **Dimensions** | **Provider** | **Language** | **Type** | | -------------------------------- | -------------- | ------------ | ------------ | --------- | | embed-english-v3 | 512 | Bedrock | English | Embedding | | embed-multilingual-v3 | 512 | Bedrock | Multilingual | Embedding | | embed-english-v3.0 | - | EdenAI | English | Embedding | | embed-english-light-v3.0 | - | EdenAI | English | Embedding | | embed-multilingual-v3.0 | - | EdenAI | Multilingual | Embedding | | embed-english-v2.0 | - | EdenAI | English | Embedding | | embed-english-light-v2.0 | - | EdenAI | English | Embedding | | embed-multilingual-v2.0 | - | EdenAI | Multilingual | Embedding | | 4096embed-english-v2.0 | 4,096 | EdenAI | English | Embedding | | 1024embed-english-light-v2.0 | 1,024 | EdenAI | English | Embedding | | 768\_\_embed-multilingual-v2.0 | 768 | EdenAI | Multilingual | Embedding | | 4096\_\_embed-english-v2.0 | 4,096 | EdenAI | English | Embedding | | 1024\_\_embed-english-light-v2.0 | 1,024 | EdenAI | English | Embedding | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a model from the **Cohere Series**, such as `command`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "cohere", "model": "command", "max_tokens": 150, "messages": [ { "role": "user", "content": "Can you explain how photosynthesis works?" } ] }' ``` --- ### Response Example The expected response structure for the Cohere model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "cohere", "model": "command", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Photosynthesis is the process by which green plants, algae, and some bacteria convert light energy into chemical energy. Here's how it works:\n\n1. **Light Absorption**: Plants capture light energy using a pigment called chlorophyll, which is found in chloroplasts.\n\n2. **Water and Carbon Dioxide**: They absorb water through their roots and carbon dioxide from the air.\n\n3. **Glucose Production**: The light energy is used to convert water and carbon dioxide into glucose (a sugar) and oxygen. The equation is:\n \n 6CO2 + 6H2O + light energy → C6H12O6 + 6O2\n\nThis process provides energy for the plant and releases oxygen into the atmosphere." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.00225, "latency_ms": 3100 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the Cohere Series, such as `command` or others suited to your task. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/features/completions). This example demonstrates how to seamlessly query models from the Cohere Series for conversational or instructional tasks. --- ### **Applications and Integrations** - **Conversational AI:** Powering chatbots, virtual assistants, and other dialogue-based systems. Try it with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI). - **Enterprise Solutions:** Leveraging Cohere's models for business applications and customer service automation. - **Content Generation:** Creating high-quality text content using [Cohere's generation capabilities](https://docs.cohere.com/docs/text-generation){rel=""nofollow""}. - **Extended Context Tasks:** Processing long documents with models supporting up to 128K tokens. Learn more in our [Models Guide](https://apipie.ai/docs/features/models). - **Multilingual Support:** Handling text processing tasks across [multiple languages](https://docs.cohere.com/docs/multilingual-language-models){rel=""nofollow""} effectively. - **RAG Applications:** Building powerful [retrieval-augmented generation systems](https://docs.cohere.com/docs/retrieval-augmented-generation-rag){rel=""nofollow""} with Cohere's models. --- ### **Ethical Considerations** Cohere models are designed with responsible AI practices in mind. Users should implement appropriate safeguards and consider potential biases in model outputs. For guidance on responsible AI usage, see Cohere's [Responsible AI Guidelines](https://cohere.com/blog/the-enterprise-guide-to-ai-safety){rel=""nofollow""}. --- ### **Licensing** Cohere models are available through their [commercial licensing terms](https://cohere.com/terms-of-use){rel=""nofollow""}. For detailed licensing information and usage terms, consult the [official Cohere documentation](https://docs.cohere.com/){rel=""nofollow""} and respective hosting providers. ::tip Try out the Cohere models in APIpie's various \[supported integrations.]\(/docs/Sandbox/Chat) :: # DeepSeek AI by APIpie: Overview & Key Features ![DeepSeek AI](https://apipie.ai/img/docs/models/DeepSeek.png){width="100%"} ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Description** The [DeepSeek Series](https://www.deepseek.com/){rel=""nofollow""} represents DeepSeek AI's family of state-of-the-art large language models. These models leverage cutting-edge technology including Mixture-of-Experts (MoE) architecture to deliver exceptional performance in natural language processing, code generation, mathematical reasoning, and extended-context interactions. The models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). #### **Key Features** - **Extended Token Capacity:** Models support context lengths from 32K to 131K tokens for handling various text processing needs. - **Multi-Provider Availability:** Accessible across platforms like [OpenRouter](https://openrouter.ai/){rel=""nofollow""}, [Together](https://www.together.ai/){rel=""nofollow""}, and [Amazon Bedrock](https://aws.amazon.com/bedrock/){rel=""nofollow""}. - **Diverse Applications:** Optimized for chat, instruction-following, analysis, writing, code generation, and mathematical reasoning. - **Mixture-of-Experts Architecture:** DeepSeek-V3 and DeepSeek-R1 utilize MoE architecture with 671B total parameters and 37B activated parameters for efficient inference. - **Open Source Models:** DeepSeek offers MIT-licensed models that support commercial use, including distilled versions based on Llama and Qwen. --- ### **Model Comparison and Monitoring** When choosing between DeepSeek models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (OpenRouter, Together, Bedrock) - Latency metrics and historical performance data - Response time comparisons between different model versions **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons and pricing trends across different providers - Usage-based cost optimization recommendations **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics This monitoring system helps users: - Compare costs and pricing across different DeepSeek models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all DeepSeek models. :: ### **Model List in the DeepSeek Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::info DeepSeek-V3 is a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. DeepSeek-R1 is built on the same architecture but focuses on enhanced reasoning capabilities through reinforcement learning. The DeepSeek-R1-Distill series offers smaller, efficient models that maintain strong performance, especially on mathematical and coding tasks. :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Provider** | **Type** | | ----------------------------- | -------------- | ------------------- | ------------ | -------- | | deepseek-chat | 64,000 | 16,000 | OpenRouter | LLM | | DeepSeek-V3 | 131,072 | 131,072 | Together | LLM | | deepseek-r1-distill-llama-8b | 32,000 | 32,000 | OpenRouter | LLM | | deepseek-r1-distill-qwen-1.5b | 131,072 | 32,768 | OpenRouter | LLM | | deepseek-r1-distill-qwen-14b | 64,000 | 64,000 | OpenRouter | LLM | | deepseek-r1-distill-qwen-32b | 131,072 | 8,192 | OpenRouter | LLM | | deepseek-r1-distill-llama-70b | 131,072 | 8,192 | OpenRouter | LLM | | deepseek-r1 | 64,000 | 16,000 | OpenRouter | LLM | | r1-v1 | - | - | Bedrock | LLM | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a model from the **DeepSeek Series**, such as `deepseek-chat`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "deepseek-chat", "max_tokens": 150, "messages": [ { "role": "user", "content": "Can you explain how photosynthesis works?" } ] }' ``` --- ### Response Example The expected response structure for the DeepSeek model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "deepseek-chat", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Photosynthesis is the process by which plants convert sunlight into chemical energy. Here's how it works:\n\n1. **Light Absorption**: Plants capture sunlight using chlorophyll in their leaves\n\n2. **Water and CO2**: They take in water through roots and carbon dioxide through leaf pores\n\n3. **Chemical Reaction**: Using sunlight's energy, they convert H2O and CO2 into glucose and oxygen:\n 6CO2 + 6H2O + light → C6H12O6 + 6O2\n\nThis process produces food for the plant and releases oxygen as a byproduct." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.000107, "latency_ms": 2243 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the DeepSeek Series suited to your task. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/features/completions). This example demonstrates how to seamlessly query models from the DeepSeek Series for conversational or instructional tasks. --- ### **DeepSeek Model Variants** #### **[DeepSeek-V3](https://github.com/deepseek-ai/DeepSeek-V3){rel=""nofollow""}** DeepSeek-V3 is a powerful Mixture-of-Experts (MoE) language model with 671B total parameters and 37B activated for each token. Pre-trained on 14.8 trillion diverse and high-quality tokens, it outperforms other open-source models and achieves performance comparable to leading closed-source models. **Key Features:** - Innovative load balancing strategy without auxiliary loss - Multi-Token Prediction (MTP) training objective - FP8 mixed precision training framework - 128K context window support - Exceptional performance on math, code, and reasoning tasks #### **DeepSeek-R1** DeepSeek-R1 is built on the same architecture as DeepSeek-V3 but focuses on enhanced reasoning capabilities through reinforcement learning. It achieves performance comparable to leading models on math, code, and reasoning tasks. **Key Features:** - Trained via large-scale reinforcement learning - Naturally emerged reasoning behaviors including self-verification and reflection - Strong performance on mathematical reasoning and coding tasks - MIT license allowing commercial use and distillation #### **[DeepSeek-R1-Distill](https://github.com/deepseek-ai/DeepSeek-R1){rel=""nofollow""} Series** The DeepSeek-R1-Distill series offers smaller, efficient models that maintain strong performance, especially on mathematical and coding tasks. These models are distilled from DeepSeek-R1 and based on Llama and Qwen architectures. **Available Models:** - DeepSeek-R1-Distill-Qwen-1.5B - DeepSeek-R1-Distill-Qwen-7B - DeepSeek-R1-Distill-Llama-8B - DeepSeek-R1-Distill-Qwen-14B - DeepSeek-R1-Distill-Qwen-32B - DeepSeek-R1-Distill-Llama-70B --- ### **Applications and Integrations** - **Conversational AI:** Powering chatbots, virtual assistants, and other dialogue-based systems. Try it with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI). - **Code Generation:** Leveraging DeepSeek's strong coding capabilities for software development, debugging, and code explanation. - **Mathematical Reasoning:** Using DeepSeek-R1 and its distilled models for complex mathematical problem-solving and education. - **Content Creation:** Using DeepSeek's writing abilities for content generation and editing. - **Extended Context Tasks:** Processing long documents with models supporting up to 131K tokens. Learn more in our [Models Guide](https://apipie.ai/docs/features/models). - **Educational Support:** Providing detailed explanations and tutoring across various subjects, especially in STEM fields. --- ### **Ethical Considerations** DeepSeek models are built with responsible AI principles in mind. Users should implement appropriate safeguards and consider potential biases in model outputs. For guidance on responsible AI usage, refer to DeepSeek's documentation and best practices. --- ### **Licensing** The DeepSeek Series is available through various API platforms. DeepSeek-V3, DeepSeek-R1, and the DeepSeek-R1-Distill series are licensed under the MIT License, which supports commercial use and allows for modifications and derivative works, including distillation for training other LLMs. For detailed licensing information, consult [DeepSeek's documentation](https://huggingface.co/deepseek-ai){rel=""nofollow""} and respective hosting providers. ::tip Try out the DeepSeek models in APIpie's various \[supported integrations.]\(/docs/Sandbox/Chat) :: # Qwen API Overview: Unlock Conversational AI ![Qwen](https://apipie.ai/img/docs/models/Qwen.jpeg){width="100%"} ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Model Comparison and Monitoring** When choosing between Qwen models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (OpenRouter, EdenAI, Together, Bedrock) - Latency metrics and historical performance data - Response time comparisons between different model versions - Specific performance tracking for Chat, Instruction, and Vision-Language variants **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different providers and model variants - Usage-based optimization recommendations - Provider-specific pricing tracking **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different capabilities (chat, instruction, vision) This monitoring system helps users: - Compare costs and pricing across different Qwen models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose optimal model based on performance and cost needs ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all Qwen models. :: ### **Description** The [Qwen Series](https://github.com/QwenLM/Qwen){rel=""nofollow""} represents a comprehensive family of transformer-based models optimized for a wide range of NLP applications. Developed by [Alibaba Cloud](https://www.alibabacloud.com/){rel=""nofollow""}, these models leverage cutting-edge technology to deliver exceptional performance in conversational AI, instruction-following tasks, and extended-context interactions. The models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). #### **Key Features** - **Extended Token Capacity:** All models support up to 32,768 tokens for efficient handling of long-text inputs and context-rich conversations. - **Multi-Provider Availability:** Accessible across platforms like [OpenRouter](https://openrouter.ai/){rel=""nofollow""}, [EdenAI](https://www.edenai.co/){rel=""nofollow""}, [Together](https://www.together.ai/){rel=""nofollow""}, and [Amazon Bedrock](https://aws.amazon.com/bedrock/){rel=""nofollow""}. - **Diverse Subtypes:** Includes Chat, Instruction, and Vision-Language variants tailored for specific applications. - **Scalability:** Models ranging from lightweight solutions (1.5B parameters) to high-capacity configurations (72B parameters) for advanced tasks. --- ### **Model List in the Qwen Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::info For information about model performance and benchmarks, see the \[Qwen Technical Report]\(https\://github.com/QwenLM/Qwen/blob/main/QWEN\_TECHNICAL\_REPORT.pdf). :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Providers** | **Subtype** | | ---------------------- | -------------- | ------------------- | ------------------------------------- | --------------- | | qwen-2-72b-instruct | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Instruction | | qwen-2-7b-instruct | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Instruction | | Qwen2-1.5B-Instruct | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Instruction | | Qwen2-7B-Instruct | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Instruction | | Qwen1.5-14B-Chat | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-1.8B-Chat | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-32B-Chat | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-7B-Chat | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-0.5B-Chat | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-4B-Chat | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | qwen-2-vl-7b-instruct | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Vision-Language | | qwen-2-vl-72b-instruct | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Vision-Language | | qwen-2.5-72b-instruct | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Instruction | | qwen-2.5-7b-instruct | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Instruction | | eva-qwen-2.5-32b | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Instruction | | Qwen2-72B-Instruct | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Instruction | | Qwen2-72B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-0.5B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-1.8B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-4B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-7B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-72B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen2-7B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen2-1.5B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-32B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | Qwen1.5-14B | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | qwen1-5\_32k | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | | qwen2\_32k | 32,768 | 32,768 | OpenRouter, EdenAI, Together, Bedrock | Chat | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a model from the **Qwen Series**, such as `qwen-2-7b-instruct`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "qwen-2-7b-instruct", "max_tokens": 150, "messages": [ { "role": "user", "content": "Can you explain how photosynthesis works?" } ] }' ``` --- ### Response Example The expected response structure for the Qwen model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "qwen-2-7b-instruct", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Photosynthesis is the process by which green plants, algae, and some bacteria convert light energy into chemical energy. Here’s how it works:\n\n1. **Light Absorption**: Plants capture light energy using a pigment called chlorophyll, which is found in chloroplasts.\n\n2. **Water and Carbon Dioxide**: They absorb water through their roots and carbon dioxide from the air.\n\n3. **Glucose Production**: The light energy is used to convert water and carbon dioxide into glucose (a sugar) and oxygen. The equation is:\n \n 6CO2 + 6H2O + light energy → C6H12O6 + 6O2\n\nThis process provides energy for the plant and releases oxygen into the atmosphere." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.00225, "latency_ms": 3100 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the Qwen Series, such as `qwen-2-7b-instruct` or others suited to your task. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/features/completions). This example demonstrates how to seamlessly query models from the Qwen Series for conversational or instructional tasks. --- ### **Applications and Integrations** - **Conversational AI:** Powering chatbots, virtual assistants, and other dialogue-based systems. Try it with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI). - **Instructional Scenarios:** Tailored for executing complex, multi-step tasks based on user inputs. See [Qwen's instruction guide](https://github.com/QwenLM/Qwen/tree/main/examples){rel=""nofollow""}. - **Vision-Language Models:** Addressing multimodal tasks combining textual and visual inputs using specialized VL models. Learn more in [Qwen-VL documentation](https://github.com/QwenLM/Qwen-VL){rel=""nofollow""}. - **Extended Context Tasks:** Providing coherent responses for long-sequence inputs. See our [Models Guide](https://apipie.ai/docs/features/models) for details. --- ### **Ethical Considerations** The Qwen models are highly capable but lack inherent moderation systems. Users are encouraged to implement safeguards for appropriate deployment in sensitive contexts. For guidance on responsible AI usage, see [Alibaba Cloud's Best Practices.](https://resource.alibabacloud.com/bestpractice){rel=""nofollow""} --- ### **Licensing** The Qwen Series is released under flexible licensing terms, allowing for both commercial and non-commercial usage. For detailed terms, consult the [official Qwen repository](https://github.com/QwenLM/Qwen/blob/main/LICENSE){rel=""nofollow""}, [model cards](https://huggingface.co/Qwen){rel=""nofollow""}, and respective hosting providers. ::tip Try out the Qwen models in APIpie's various \[supported integrations.]\(/docs/Sandbox/Chat) :: # Perplexity API Overview: Key Features & Use Cases ![Perplexity](https://apipie.ai/img/docs/models/Perplexity.png) ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/1.Features/10.Models) and [Completions Guide](https://apipie.ai/docs/1.Features/15.Completions). :: ### **Model Comparison and Monitoring** When choosing between Perplexity models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (OpenRouter) - Latency metrics and historical performance data - Response time comparisons between different model sizes - Specific performance tracking for online inference capabilities **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/1.Features/10.Models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different Sonar variants - Usage-based optimization recommendations - Real-time inference cost tracking **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different model sizes (small, large, huge) This monitoring system helps users: - Compare costs and pricing across different Perplexity models - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose optimal model size based on performance and cost needs ::tip Use the [Models Route](https://apipie.ai/docs/1.Features/10.Models) to access real-time pricing and performance data for all Perplexity models. :: ### **Description** [Perplexity AI](https://docs.perplexity.ai/guides/model-cards){rel=""nofollow""} is pioneering the development of advanced language models with a focus on real-time information processing and extended context understanding. Building upon Meta's [Llama architecture](https://ai.meta.com/llama/){rel=""nofollow""}, Perplexity has enhanced these foundation models with specialized training and optimizations for online inference. Their Sonar models represent a significant breakthrough, combining Llama's powerful language understanding capabilities with real-time information processing and extended context windows. This innovative approach leverages Llama's strong foundation while adding crucial features for real-world applications: - **Online Inference Optimization:** Enhanced the base Llama architecture for real-time processing - **Extended Context Understanding:** Improved context window handling up to 128K tokens - **Knowledge Integration:** Added capabilities for real-time information access and processing These models are available through [APIpie's routing system](https://apipie.ai/docs/1.Features/35.Routing), offering state-of-the-art performance for various applications. #### **Key Features** - **Extended Context Processing:** Supports up to 128K tokens context length for comprehensive document analysis and long-form conversations - **Online Inference:** Designed for real-time information processing and up-to-date knowledge - **Scalable Architecture:** Available in multiple sizes (small, large, huge) to suit different computational needs - **High Performance:** Optimized for both speed and accuracy in natural language understanding tasks - **Knowledge Integration:** Enhanced with real-time information access capabilities - **Adaptive Learning:** Continuously updated to maintain current knowledge and capabilities --- ### **Available Models** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/1.Features/10.Models#fetching-models) for the up to date list of models** ::note Perplexity offers various models optimized for different use cases. The following models represent their Sonar series, which excels at real-time information processing with extended context windows. For detailed information about model capabilities and use cases, visit the [Perplexity Blog](https://blog.perplexity.ai/){rel=""nofollow""}. :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Providers** | **Subtype** | | --------------------------------- | -------------- | ------------------- | ------------- | ----------- | | llama-3.1-sonar-small-128k-online | 127,072 | 127,072 | OpenRouter | Chat | | llama-3.1-sonar-large-128k-online | 127,072 | 127,072 | OpenRouter | Chat | | llama-3.1-sonar-huge-128k-online | 127,072 | 127,072 | OpenRouter | Chat | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/Features/Completions) to interact with a Perplexity model: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "llama-3.1-sonar-huge-128k-online", "max_tokens": 150, "messages": [ { "role": "user", "content": "What are the latest developments in quantum computing?" } ] }' ``` --- ### Response Example The expected response structure for a Perplexity model: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "llama-3.1-sonar-huge-128k-online", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Recent developments in quantum computing include advances in error correction, new qubit architectures, and breakthrough experiments in quantum supremacy. Companies like IBM, Google, and others continue to make progress in both hardware and software aspects of quantum computing technology." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.00225, "latency_ms": 3100 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify a provider or leave blank for [automatic selection](https://apipie.ai/docs/1.Features/35.Routing) - **Model:** Choose from available Perplexity models based on your needs. See [Models Guide](https://apipie.ai/docs/1.Features/10.Models#fetching-models) - **Max Tokens:** Leverage the extended context window of up to 127K tokens - **Messages:** Format your request with user inputs and system instructions. See [message formatting](https://apipie.ai/docs/1.Features/15.Completions) --- ### **Applications and Use Cases** - **Real-time Information Processing:** Ideal for applications requiring current information and knowledge - **Long-form Content Analysis:** Perfect for processing and analyzing lengthy documents, research papers, or conversations - **Knowledge-intensive Tasks:** Excellent for research, analysis, and complex problem-solving requiring deep understanding - **Enterprise Solutions:** Suitable for business applications requiring processing of large documents and real-time information - **Educational Applications:** Valuable for in-depth learning and research assistance - **Research and Development:** Powerful tools for scientific research and technological innovation - **Content Generation:** Advanced capabilities for creating high-quality, well-researched content Try these models with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI) for an interactive experience. --- ### **Best Practices** - Utilize the extended context window effectively by providing comprehensive context when needed - Consider the model size based on your specific use case - smaller models for faster responses, larger models for more complex tasks - Implement appropriate error handling and retry mechanisms for optimal performance - Follow [Perplexity's usage guidelines](https://www.perplexity.ai/hub/getting-started){rel=""nofollow""} for best results - Leverage the real-time information processing capabilities for up-to-date responses - Structure prompts to take advantage of the models' knowledge integration features --- ### **Resources and Documentation** - [Perplexity AI Documentation](https://docs.perplexity.ai/){rel=""nofollow""} - [APIpie Models Guide](https://apipie.ai/docs/1.Features/10.Models) - [Perplexity Research Blog](https://blog.perplexity.ai/){rel=""nofollow""} - [Model Performance Benchmarks](https://blog.perplexity.ai/blog/introducing-pplx-online-llms){rel=""nofollow""} ::tip Experience the power of Perplexity models in APIpie's various [supported integrations](https://apipie.ai/docs/10.Integrations). :: # Explore Google's AI: Gemini & Gemma Series Overview ![Google Gemini](https://apipie.ai/img/docs/models/Gemini.png) ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Description** [Google's AI models](https://ai.google/get-started/our-models/){rel=""nofollow""}, including the [Gemini Series](https://blog.google/technology/ai/google-gemini-ai/){rel=""nofollow""} and [Gemma Series](https://ai.google.dev/gemma/){rel=""nofollow""}, represent cutting-edge advancements in artificial intelligence. Developed by [Google DeepMind](https://deepmind.google/){rel=""nofollow""}, these models leverage state-of-the-art technology to deliver exceptional performance across various AI tasks. The models are available through multiple providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). The [Gemini Series](https://deepmind.google/technologies/gemini/){rel=""nofollow""} represents Google's most advanced AI models, capable of sophisticated reasoning across text, code, images, and video. Meanwhile, the [Gemma Series](https://ai.google.dev/gemma){rel=""nofollow""} offers efficient, open-source models built on the same research and technology as Gemini. These models are accessible through [Google Cloud's Vertex AI platform](https://cloud.google.com/vertex-ai/docs/generative-ai/learn/models){rel=""nofollow""} and various third-party providers. #### **Key Features** - **Advanced Token Processing:** Models support context lengths from 4K to 2M tokens for comprehensive text processing, with [optimized performance](https://cloud.google.com/vertex-ai/docs/generative-ai/learn/models#context_window){rel=""nofollow""} across different scales. - **Multi-Provider Access:** Available through platforms like [Google Cloud](https://cloud.google.com/vertex-ai){rel=""nofollow""}, [OpenRouter](https://openrouter.ai/){rel=""nofollow""}, [EdenAI](https://www.edenai.co/){rel=""nofollow""}, and more, with [enterprise-grade security](https://cloud.google.com/security/products){rel=""nofollow""}. - **Versatile Applications:** Optimized for chat, instruction-following, code generation, and [multimodal tasks](https://deepmind.google/technologies/gemini/#capabilities){rel=""nofollow""}, with state-of-the-art performance across domains. --- ### **Model Comparison and Monitoring** When choosing between Google's models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (Together, Deepinfra, OpenRouter, EdenAI) - Latency metrics and historical performance data - Response time comparisons between different model versions and families - Specific performance tracking for Gemini and Gemma variants **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different providers and model types - Usage-based optimization recommendations - Separate tracking for proprietary and open-source model costs **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different model capabilities (text, code, vision) This monitoring system helps users: - Compare costs and pricing across different Google models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose between proprietary and open-source options based on needs ::tip Use the [Models Route](https://apipie.ai/docs/features/models) to access real-time pricing and performance data for all Google models. :: ### **Model List in the Google Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/1.Features/10.Models#fetching-models) for the up to date list of models** ::note For information about model performance and benchmarks, see the [Gemini Technical Report](https://storage.googleapis.com/deepmind-media/gemini/gemini_1_report.pdf){rel=""nofollow""} and [Gemma Technical Report](https://storage.googleapis.com/deepmind-media/gemma/gemma-report.pdf){rel=""nofollow""}. :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Providers** | **Subtype** | | ------------------------- | -------------- | ------------------- | ----------------------------- | ----------- | | gemini-pro-1.5 | 2,000,000 | 8,192 | OpenRouter | Text | | gemini-flash-1.5 | 1,000,000 | 8,192 | OpenRouter | Chat | | gemini-flash-1.5-8b | 1,000,000 | 8,192 | OpenRouter | Chat | | gemini-pro | 91,728 | 22,937 | EdenAI | Chat | | palm-2-codechat-bison-32k | 91,750 | 22,937 | OpenRouter | Code | | palm-2-chat-bison-32k | 32,768 | 8,192 | OpenRouter | Chat | | gemini-pro | 32,760 | 8,192 | OpenRouter | Text | | palm-2-codechat-bison | 20,070 | 2,867 | OpenRouter | Code | | palm-2-chat-bison | 9,216 | 1,024 | OpenRouter | Chat | | gemma-7b-it | 8,192 | 8,192 | Deepinfra | Chat | | gemma-2-27b-it | 8,192 | 8,192 | Together | Chat | | gemma-2b-it | 8,192 | 8,192 | Together | Chat | | gemma-2-27b-it | 8,192 | 4,096 | OpenRouter | Text | | gemma-2-9b-it | 4,096 | 4,096 | OpenRouter, Monster, Together | Chat | | gemini-pro-vision | 16,384 | 2,048 | OpenRouter, EdenAI | Vision | | gemini-1.5-flash | - | - | EdenAI | Chat | | gemini-1.5-pro | - | - | EdenAI | Chat | | gemini-1.5-flash-latest | - | - | EdenAI | Chat | | gemini-1.5-pro-exp-0801 | - | - | EdenAI | Chat | | gemini-1.5-pro-exp-0827 | - | - | EdenAI | Chat | | gemini-1.5-pro-latest | - | - | EdenAI | Chat | | gemini-1.5-flash-8b | - | - | EdenAI | Chat | | chat-bison | - | - | EdenAI | Chat | | gemma-1.1-7b-it | - | - | Deepinfra | Chat | | vit-base-patch16-224 | - | - | Deepinfra | Image | | vit-base-patch16-384 | - | - | Deepinfra | Image | | textembedding-gecko | 768 | - | EdenAI | Embedding | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a model from the **Google Series**, such as `gemini-pro-1.5`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "gemini-pro-1.5", "max_tokens": 150, "messages": [ { "role": "user", "content": "What are the key differences between machine learning and deep learning?" } ] }' ``` --- ### Response Example The expected response structure for a Google model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "gemini-pro-1.5", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Machine learning and deep learning differ in several key aspects:\n\n1. **Complexity**: Machine learning uses simpler algorithms for pattern recognition, while deep learning uses complex neural networks with multiple layers.\n\n2. **Data Requirements**: Machine learning can work with smaller datasets, but deep learning typically needs vast amounts of data to be effective.\n\n3. **Feature Extraction**: Machine learning often requires manual feature engineering, while deep learning automatically learns and extracts relevant features.\n\n4. **Hardware Requirements**: Machine learning algorithms can run on standard computers, but deep learning usually needs powerful GPUs for efficient processing.\n\n5. **Applications**: Machine learning is suited for structured data analysis, while deep learning excels in complex tasks like image recognition and natural language processing." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 18, "completion_tokens": 132, "total_tokens": 150, "prompt_characters": 72, "response_characters": 548, "cost": 0.003, "latency_ms": 2800 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the Google Series, such as `gemini-pro-1.5` or others suited to your task. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/features/completions). This example demonstrates how to seamlessly query models from the Google Series for conversational or instructional tasks. --- ### **Applications and Integrations** - **Enterprise Solutions:** Powering business applications with [Google Cloud's Vertex AI](https://cloud.google.com/vertex-ai){rel=""nofollow""}, including [industry-specific solutions](https://cloud.google.com/solutions#industry-solutions){rel=""nofollow""}. - **Multimodal Tasks:** Using vision-enabled models for image understanding and analysis with [Gemini Pro Vision](https://deepmind.google/technologies/gemini/){rel=""nofollow""}, supporting advanced [computer vision applications](https://cloud.google.com/vertex-ai/docs/generative-ai/multimodal/overview){rel=""nofollow""}. - **Research and Development:** Leveraging Gemma's open-source capabilities for innovation and experimentation, with [comprehensive documentation](https://ai.google.dev/gemma/docs){rel=""nofollow""} and [community resources](https://github.com/google/gemma_pytorch){rel=""nofollow""}. - **Extended Context Processing:** Handling long documents with models supporting up to 2M tokens, ideal for document analysis and [knowledge processing](https://cloud.google.com/vertex-ai/docs/generative-ai/learn/overview){rel=""nofollow""}. Learn more in our [Models Guide](https://apipie.ai/docs/features/models). - **Chat Applications:** Building conversational AI systems with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI), leveraging [Google's chat model best practices](https://cloud.google.com/vertex-ai/docs/generative-ai/chat/chat-prompts){rel=""nofollow""}. --- ### **Ethical Considerations** Google's AI models are designed with responsible AI principles in mind. Users should implement appropriate safeguards and consider potential biases in model outputs. For guidance on responsible AI usage, see Google's [AI Principles](https://ai.google/responsibility/principles/){rel=""nofollow""}, [Responsible AI Practices](https://ai.google/responsibility/responsible-ai-practices/){rel=""nofollow""}. --- ### **Licensing** The Google Series includes both proprietary and open-source models. While Gemini models are available through Google Cloud with specific terms of service, [Gemma models](https://github.com/google-deepmind/gemma/blob/main/LICENSE){rel=""nofollow""} are released under the Apache 2.0 license for research and commercial use. For detailed licensing information, consult the [model-specific documentation](https://cloud.google.com/vertex-ai/docs/generative-ai/learn/model-versioning){rel=""nofollow""}, usage guidelines. ::tip Try out the Google models in APIpie's various [supported integrations.](https://apipie.ai/docs/Sandbox/Chat) :: # Grok API Overview: Unlock Real-Time AI Solutions ![Grok](https://apipie.ai/img/docs/models/Grok.png) ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/1.Features/10.Models) and [Completions Guide](https://apipie.ai/docs/1.Features/15.Completions). :: ### **Description** The [Grok Series](https://x.ai/){rel=""nofollow""} represents xAI's family of state-of-the-art large language models. These models, developed by [xAI](https://x.ai/){rel=""nofollow""}, leverage cutting-edge technology to deliver exceptional performance in natural language processing, instruction-following tasks, and multimodal interactions. Grok is known for its real-time knowledge capabilities through [X platform integration](https://twitter.com/grok){rel=""nofollow""} and its unique personality that combines intelligence with wit. The models are available through [APIpie's routing system](https://apipie.ai/docs/features/routing). For technical details and performance metrics, see the [Grok technical report](https://x.ai/blog/grok-os){rel=""nofollow""}. #### **Key Features** - **Extended Token Capacity:** Models support context lengths up to 131K tokens for handling extensive text processing needs. - **Multimodal Capabilities:** Vision-enabled models for processing both text and images with state-of-the-art performance. - **Real-time Knowledge:** Access to current information through X platform integration, setting it apart from other LLMs. - **Provider Availability:** Accessible through [OpenRouter](https://openrouter.ai/){rel=""nofollow""}. - **Diverse Applications:** Optimized for chat, instruction-following, vision tasks, and real-time data analysis. - **Unique Personality:** Combines factual responses with wit and humor, creating engaging interactions. --- ### **Model Comparison and Monitoring** When choosing between Grok models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking through OpenRouter - Latency metrics and historical performance data - Response time comparisons between different model versions - Specific performance tracking for vision-enabled models **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different Grok versions - Usage-based optimization recommendations - Vision model cost tracking **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different capabilities (text, vision) This monitoring system helps users: - Compare costs and pricing across different Grok models - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose between standard and vision-enabled models based on needs ::tip Use the [Models Route](https://apipie.ai/docs/1.Features/10.Models) to access real-time pricing and performance data for all Grok models. :: ### **Model List in the Grok Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/1.Features/10.Models#fetching-models) for the up to date list of models** | **Model Name** | **Max Tokens** | **Response Tokens** | **Provider** | **Subtype** | **Type** | | ------------------ | -------------- | ------------------- | ------------ | ----------- | -------- | | grok-beta | 131,072 | 131,072 | openrouter | chatx | llm | | grok-vision-beta | 8,192 | 8,192 | openrouter | multimodal | vision | | grok-2-1212 | 131,072 | 131,072 | openrouter | chatx | llm | | grok-2-vision-1212 | 32,768 | 32,768 | openrouter | multimodal | vision | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/Features/Completions) to interact with a model from the **Grok Series**, such as `grok-beta`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "grok-beta", "max_tokens": 150, "messages": [ { "role": "user", "content": "Can you explain how photosynthesis works?" } ] }' ``` --- ### Response Example The expected response structure for the Grok model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "grok-beta", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Photosynthesis is the process by which green plants, algae, and some bacteria convert light energy into chemical energy. Here's how it works:\n\n1. **Light Absorption**: Plants capture light energy using a pigment called chlorophyll, which is found in chloroplasts.\n\n2. **Water and Carbon Dioxide**: They absorb water through their roots and carbon dioxide from the air.\n\n3. **Glucose Production**: The light energy is used to convert water and carbon dioxide into glucose (a sugar) and oxygen. The equation is:\n \n 6CO2 + 6H2O + light energy → C6H12O6 + 6O2\n\nThis process provides energy for the plant and releases oxygen into the atmosphere." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.00225, "latency_ms": 3100 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify "openrouter" as the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the Grok Series, such as `grok-beta` or `grok-vision-beta`. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/Features/Completions). This example demonstrates how to seamlessly query models from the Grok Series for conversational or instructional tasks. --- ### **Applications and Integrations** - **Conversational AI:** Powering chatbots, virtual assistants, and other dialogue-based systems. Try it with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI). - **Vision Tasks:** Using vision-enabled models like `grok-vision-beta` and `grok-2-vision-1212` for image understanding and multimodal applications. Learn more about [vision capabilities](https://x.ai/blog/grok-1.5v){rel=""nofollow""}. - **Extended Context Tasks:** Processing long documents with models supporting up to 131K tokens. Learn more in our [Models Guide](https://apipie.ai/docs/features/models). - **Real-time Analysis:** Leveraging Grok's connection to the X platform for current events, trends, and social media analysis. - **Research & Development:** Supporting academic and industrial research with reproducible results --- ### **Ethical Considerations** The Grok models are powerful tools that should be used responsibly. Users should implement appropriate safeguards and consider potential biases in model outputs. For guidance on responsible AI usage, refer to [AI Safety Approach](https://www.safe.ai/){rel=""nofollow""} and [Ethical AI Guidelines](https://www.unesco.org/en/artificial-intelligence/recommendation-ethics){rel=""nofollow""}. The development team maintains transparency through their open science initiative and regular safety assessments. --- ### **Licensing** For detailed licensing information and usage terms, consult [xAI terms](https://x.ai/legal/terms-of-service){rel=""nofollow""}. ::tip Try out the Grok models in APIpie's various [supported integrations.](https://apipie.ai/docs/Sandbox/Chat) :: # AI21 Labs: Powerful AI & API Integration Guide ![AI21 Labs](https://apipie.ai/img/docs/models/AI21.png){width="100%"} ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Description** [AI21 Labs](https://www.ai21.com/){rel=""nofollow""} is a leading AI research company specializing in advanced language models. Their [Jurassic-2](https://www.ai21.com/blog/introducing-j2/){rel=""nofollow""} (J2) and [Jamba](https://www.ai21.com/jamba){rel=""nofollow""} series models leverage state-of-the-art technology to deliver exceptional performance in natural language processing and instruction-following tasks. The models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). The Jurassic-2 series offers powerful base models optimized for enterprise use, while the Jamba series represents their latest advancement in large language models, featuring extended context windows and enhanced instruction-following capabilities. Both series are built on AI21's proprietary training techniques and architectures detailed in their [technical documentation](https://docs.ai21.com/docs/overview){rel=""nofollow""}. #### **Key Features** - **Extended Token Capacity:** Models support context lengths from 8K to 256K tokens for handling various text processing needs. - **Multi-Provider Availability:** Accessible through platforms like [Amazon Bedrock](https://aws.amazon.com/bedrock/){rel=""nofollow""} and [OpenRouter](https://openrouter.ai/){rel=""nofollow""}. - **Diverse Applications:** Optimized for chat, instruction-following, and general text generation tasks. - **Enterprise Focus:** Models designed for production-ready deployment with consistent performance and reliability. --- ### **Model List in the AI21 Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::note For information about model performance and benchmarks, see the [AI21 Labs Documentation](https://docs.ai21.com/){rel=""nofollow""}. :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Providers** | **Subtype** | **Type** | | ------------------ | -------------- | ------------------- | ------------- | ----------- | -------- | | j2-grande-instruct | 8,192 | 8,192 | Bedrock | LLM | LLM | | j2-mid | 8,000 | 8,191 | Bedrock | LLM | LLM | | j2-mid-v1 | 8,192 | 8,192 | Bedrock | LLM | LLM | | j2-ultra-v1 | 8,192 | 8,192 | Bedrock | LLM | LLM | | jamba-instruct-v1 | - | - | Bedrock | LLM | LLM | | jamba-instruct | 256,000 | 4,096 | OpenRouter | ChatX | LLM | | jamba-1-5-large | 256,000 | 4,096 | OpenRouter | ChatX | LLM | | jamba-1-5-mini | 256,000 | 4,096 | OpenRouter | ChatX | LLM | | jamba-1-5-large-v1 | 256,000 | 4,096 | Bedrock | LLM | LLM | | jamba-1-5-mini-v1 | 256,000 | 4,096 | Bedrock | LLM | LLM | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a model from the **AI21 Series**, such as `jamba-instruct`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "jamba-instruct", "max_tokens": 150, "messages": [ { "role": "user", "content": "Can you explain how photosynthesis works?" } ] }' ``` --- ### Response Example The expected response structure for the AI21 model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "jamba-instruct", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Photosynthesis is the process by which green plants, algae, and some bacteria convert light energy into chemical energy. Here's how it works:\n\n1. **Light Absorption**: Plants capture light energy using a pigment called chlorophyll, which is found in chloroplasts.\n\n2. **Water and Carbon Dioxide**: They absorb water through their roots and carbon dioxide from the air.\n\n3. **Glucose Production**: The light energy is used to convert water and carbon dioxide into glucose (a sugar) and oxygen. The equation is:\n \n 6CO2 + 6H2O + light energy → C6H12O6 + 6O2\n\nThis process provides energy for the plant and releases oxygen into the atmosphere." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.00225, "latency_ms": 3100 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the AI21 Series, such as `jamba-instruct` or others suited to your task. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/features/completions). This example demonstrates how to seamlessly query models from the AI21 Series for conversational or instructional tasks. --- ### **Model Comparison and Monitoring** When choosing between AI21 models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (Bedrock, OpenRouter) - Latency metrics and historical performance data - Response time comparisons between different model versions **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons and pricing trends across different providers **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics This monitoring system helps users: - Compare costs and pricing across different AI21 models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues ::note Use the [Models Route](https://apipie.ai/docs/features/models) to access real-time pricing and performance data for all AI21 models. :: ### **Applications and Integrations** - **Conversational AI:** Powering chatbots, virtual assistants, and other dialogue-based systems. Try it with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI). - **Enterprise Solutions:** Leveraging [AI21 Studio](https://ai21.com/){rel=""nofollow""} for customized AI applications and workflows. - **Content Generation:** Creating high-quality content with models optimized for natural language generation. - **Extended Context Tasks:** Processing long documents with models supporting up to 256K tokens. Learn more in our [Models Guide](https://apipie.ai/docs/features/models). --- ### **Ethical Considerations** The AI21 models are powerful tools that should be used responsibly. Users should implement appropriate safeguards and consider potential biases in model outputs. For guidance on responsible AI usage, see AI21's [Responsible use Guidelines](https://docs.ai21.com/docs/responsible-use-1){rel=""nofollow""}. --- ### **Licensing** AI21 models are available through commercial licensing agreements. For detailed licensing information and terms of use, consult the [AI21 Labs Terms of Service.](https://www.ai21.com/terms-of-service){rel=""nofollow""}. ::note Try out the AI21 models in APIpie's various [supported integrations.](https://apipie.ai/docs/Sandbox/Chat) :: # Hermes Overview: Key Features & Integration Guide ::div{.docs-image-row} ![Hermes](https://apipie.ai/img/docs/models/Hermes.png){width="100%"} :: ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/Features/Completions). :: ### **Model Comparison and Monitoring** When choosing between Hermes models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (OpenRouter, Together, Deepinfra) - Latency metrics and historical performance data - Response time comparisons between different model versions - Specific performance tracking for vision-enabled models **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different providers and model variants - Usage-based optimization recommendations - Provider-specific pricing tracking **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different capabilities (text, vision) This monitoring system helps users: - Compare costs and pricing across different Hermes models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose optimal provider based on performance and cost needs ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all Hermes models. :: ### **Description** The Hermes Series represents a collection of advanced large language models, primarily developed by [Nous Research](https://nousresearch.com/){rel=""nofollow""} and the open-source AI community. These models are fine-tuned versions built on top of powerful base models like [Llama](https://ai.meta.com/llama/){rel=""nofollow""}, [Mistral](https://mistral.ai/){rel=""nofollow""}, and [Mixtral](https://mistral.ai/news/mixtral-of-experts/){rel=""nofollow""}. The models leverage state-of-the-art [Direct Preference Optimization (DPO)](https://arxiv.org/abs/2305.18290){rel=""nofollow""} and other advanced training techniques for enhanced performance in natural language processing and instruction-following tasks. You can explore the models' technical details and training methodologies on their [Hugging Face model cards](https://huggingface.co/NousResearch){rel=""nofollow""}. The models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). #### **Key Features** - **Extended Token Capacity:** Models support context lengths from 4K to 32K tokens for handling various text processing needs, with advanced [attention mechanisms](https://arxiv.org/abs/1706.03762){rel=""nofollow""} for efficient processing. - **Multi-Provider Availability:** Accessible across platforms like [OpenRouter](https://openrouter.ai/){rel=""nofollow""}, [Together](https://www.together.ai/){rel=""nofollow""}, and [Deepinfra](https://deepinfra.com/){rel=""nofollow""}. - **Diverse Applications:** Optimized for chat, instruction-following, and vision tasks through specialized [multi-task training](https://arxiv.org/abs/2302.06646){rel=""nofollow""}. - **Advanced Architecture:** Built on state-of-the-art base models with specialized fine-tuning using techniques like [constitutional AI](https://arxiv.org/abs/2212.08073){rel=""nofollow""} and [RLHF](https://arxiv.org/abs/2203.02155){rel=""nofollow""}. --- ### **Model List in the Hermes Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** | **Model Name** | **Max Tokens** | **Response Tokens** | **Provider** | **Type** | | ------------------------------ | -------------- | ------------------- | ------------ | -------- | | nous-hermes-llama2-13b | 4,096 | 4,096 | openrouter | llm | | nous-hermes-2-mixtral-8x7b-dpo | 32,768 | 32,768 | openrouter | llm | | nous-hermes-2-vision-7b | 4,096 | 4,096 | openrouter | vision | | openhermes-2.5-mistral-7b | 4,096 | 4,096 | openrouter | llm | | chronos-hermes-13b-v2 | 4,096 | 4,096 | deepinfra | llm | | Nous-Hermes-2-Mixtral-8x7B-DPO | 32,768 | 32,768 | together | llm | ::note For information about model performance and benchmarks, see the [Nous Research Blog](https://nousresearch.com/blog){rel=""nofollow""} and explore detailed evaluations on the [Open LLM Leaderboard](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard){rel=""nofollow""}. :: --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/Features/Completions) to interact with a model from the **Hermes Series**, such as `nous-hermes-2-mixtral-8x7b-dpo`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "nous-hermes-2-mixtral-8x7b-dpo", "max_tokens": 150, "messages": [ { "role": "user", "content": "Can you explain how photosynthesis works?" } ] }' ``` --- ### Response Example The expected response structure for the Hermes model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "nous-hermes-2-mixtral-8x7b-dpo", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Photosynthesis is the process by which green plants, algae, and some bacteria convert light energy into chemical energy. Here's how it works:\n\n1. **Light Absorption**: Plants capture light energy using a pigment called chlorophyll, which is found in chloroplasts.\n\n2. **Water and Carbon Dioxide**: They absorb water through their roots and carbon dioxide from the air.\n\n3. **Glucose Production**: The light energy is used to convert water and carbon dioxide into glucose (a sugar) and oxygen. The equation is:\n \n 6CO2 + 6H2O + light energy → C6H12O6 + 6O2\n\nThis process provides energy for the plant and releases oxygen into the atmosphere." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.00225, "latency_ms": 3100 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the Hermes Series, such as `nous-hermes-2-mixtral-8x7b-dpo` or others suited to your task. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/Features/Completions). This example demonstrates how to seamlessly query models from the Hermes Series for conversational or instructional tasks. --- ### **Applications and Integrations** - **Conversational AI:** Powering chatbots, virtual assistants, and other dialogue-based systems. Try it with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat), [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI), or integrate with popular frameworks like [LangChain](https://python.langchain.com/){rel=""nofollow""} and [LlamaIndex](https://www.llamaindex.ai/){rel=""nofollow""}. - **Vision Tasks:** Using vision-enabled models like `nous-hermes-2-vision-7b` for image understanding and multimodal applications, leveraging [advanced vision-language architectures](https://arxiv.org/abs/2303.08891){rel=""nofollow""}. - **Extended Context Tasks:** Processing longer documents with models supporting up to 32K tokens, utilizing efficient [attention mechanisms](https://arxiv.org/abs/2307.08621){rel=""nofollow""}. Learn more in our [Models Guide](https://apipie.ai/docs/features/models). - **Instruction Following:** Leveraging advanced fine-tuning techniques like [DPO](https://arxiv.org/abs/2305.18290){rel=""nofollow""} and [Constitutional AI](https://arxiv.org/abs/2212.08073){rel=""nofollow""} for improved task completion and instruction following. - **Research and Development:** Contributing to open-source AI development through [model merging](https://arxiv.org/abs/2306.01708){rel=""nofollow""}, [knowledge distillation](https://arxiv.org/abs/2310.05344){rel=""nofollow""}, and other techniques. --- ### **Ethical Considerations** The Hermes models are powerful tools that should be used responsibly. Users should implement appropriate safeguards and consider potential biases in model outputs. --- ### **Licensing** The Hermes Series models are available under various licenses depending on their base models and providers. For detailed licensing information, consult the respective model repositories on [Hugging Face](https://huggingface.co/NousResearch){rel=""nofollow""}. ::tip Try out the Hermes models in APIpie's various \[supported integrations.]\(/docs/Sandbox/Chat) :: # Llama Models Overview: Unlock AI Potential ::div{.docs-image-row} ![Meta Llama](https://apipie.ai/img/docs/models/Meta.png){width="100%"} :: ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Model Comparison and Monitoring** When choosing between Llama models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (OpenRouter, EdenAI, Together, Deepinfra, Monster, Bedrock) - Latency metrics and historical performance data - Response time comparisons between different model versions and families - Specific performance tracking for specialized variants (Code Llama, Llama Guard) **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different providers and model variants - Usage-based optimization recommendations - Provider-specific pricing tracking **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different capabilities (text, code, vision, moderation) This monitoring system helps users: - Compare costs and pricing across different Llama models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose optimal provider based on performance and cost needs ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all Llama models. :: ### **Description** The [Llama Series](https://ai.meta.com/llama/){rel=""nofollow""} represents Meta's family of state-of-the-art large language models. These open-source models, developed by [Meta AI](https://ai.meta.com/){rel=""nofollow""}, leverage cutting-edge technology to deliver exceptional performance in natural language processing, instruction-following tasks, and extended-context interactions. The models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). #### **Key Features** - **Extended Token Capacity:** Models support context lengths from 4K to 131K tokens for handling various text processing needs. - **Multi-Provider Availability:** Accessible across platforms like [OpenRouter](https://openrouter.ai/){rel=""nofollow""}, [EdenAI](https://www.edenai.co/){rel=""nofollow""}, [Together](https://www.together.ai/){rel=""nofollow""}, [Deepinfra](https://deepinfra.com/){rel=""nofollow""}, [Monster](https://monsterapi.ai/){rel=""nofollow""}, and [Amazon Bedrock](https://aws.amazon.com/bedrock/){rel=""nofollow""}. - **Diverse Applications:** Optimized for chat, instruction-following, code generation, vision, and moderation tasks. - **Scalability:** Models ranging from efficient 1B parameter configurations to powerful 405B parameter versions. --- ### **Model List in the Llama Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::note For information about model performance and benchmarks, see the [Llama Technical Report](https://ai.meta.com/research/publications/llama-2-open-foundation-and-fine-tuned-chat-models/){rel=""nofollow""}. :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Provider** | **Type** | | -------------------------------------- | -------------- | ------------------- | ------------ | ---------- | | nous-hermes-llama2-13b | 4,096 | 4,096 | openrouter | llm | | llama-2-13b-chat | 4,096 | 4,096 | openrouter | llm | | llama2-13b-chat-v1 | 4,096 | 3,996 | bedrock | llm | | llama2-70b-chat-v1 | 4,096 | 3,996 | bedrock | llm | | Llama-2-70b-chat-hf | 4,096 | 4,096 | deepinfra | llm | | CodeLlama-70b-Instruct-hf | 4,096 | 4,096 | deepinfra | code | | CodeLlama-34b-Instruct-hf | 16,384 | 16,384 | deepinfra | code | | Phind-CodeLlama-34B-v2 | 16,384 | 16,384 | deepinfra | code | | Llama-2-7b-chat-hf | 4,096 | 4,096 | deepinfra | llm | | Llama-2-13b-chat-hf | 4,096 | 4,096 | deepinfra | llm | | Meta-Llama-3-70B-Instruct | 8,192 | 8,192 | deepinfra | llm | | Meta-Llama-3-8B-Instruct | 8,192 | 8,192 | deepinfra | llm | | llama-3-lumimaid-8b | 24,576 | 2,048 | openrouter | llm | | llama-guard-2-8b | 8,192 | 8,192 | openrouter | llm | | llama-3-lumimaid-70b | 8,192 | 2,048 | openrouter | llm | | llama3-8b-instruct-v1 | 8,192 | 8,092 | bedrock | llm | | llama3-70b-instruct-v1 | 8,192 | 8,092 | bedrock | llm | | Meta-Llama-3-8B-Instruct | 8,192 | 8,192 | monster | llm | | llama-3-8b-instruct | 8,192 | 8,192 | openrouter | llm | | llama-3-70b-instruct | 8,192 | 8,192 | openrouter | llm | | llama-3.1-sonar-large-128k-online | 127,072 | 127,072 | openrouter | llm | | llama-3.1-sonar-large-128k-chat | 131,072 | 131,072 | openrouter | llm | | llama-3.1-sonar-small-128k-online | 127,072 | 127,072 | openrouter | llm | | llama-3.1-sonar-small-128k-chat | 131,072 | 131,072 | openrouter | llm | | llama-3.1-sonar-huge-128k-online | 127,072 | 127,072 | openrouter | llm | | llama-3.1-lumimaid-8b | 32,768 | 2,048 | openrouter | llm | | meta-llama-3.1-8b-instruct | 8,192 | 8,192 | openrouter | llm | | llama-3.2-3b-instruct | 131,000 | 131,000 | openrouter | llm | | llama-3.2-1b-instruct | 131,072 | 131,072 | openrouter | llm | | llama-3.2-90b-vision-instruct | 131,072 | 131,072 | openrouter | vision | | llama-3.2-11b-vision-instruct | 131,072 | 4,096 | openrouter | vision | | llama-3.1-nemotron-70b-instruct | 131,000 | 131,000 | openrouter | llm | | llama-3.1-lumimaid-70b | 16,384 | 2,048 | openrouter | llm | | llama-3.3-70b-instruct | 131,072 | 131,072 | openrouter | llm | | llama3-1-405b-instruct-v1:0 | - | - | edenai | llm | | llama3-1-70b-instruct-v1:0 | - | - | edenai | llm | | llama3-1-8b-instruct-v1:0 | - | - | edenai | llm | | llama3-70b-instruct-v1:0 | - | - | edenai | llm | | llama3-8b-instruct-v1:0 | - | - | edenai | llm | | eva-llama-3.33-70b | 16,384 | 4,096 | openrouter | llm | | Llama-3.1-Nemotron-70B-Instruct-HF | 32,768 | 32,768 | together | llm | | Llama-3.3-70B-Instruct-Turbo | 131,072 | 131,072 | together | llm | | Llama-3.2-11B-Vision-Instruct-Turbo | 131,072 | 131,072 | together | llm | | Llama-3-8b-chat-hf | 8,192 | 8,192 | together | llm | | Llama-Guard-3-11B-Vision-Turbo | 131,072 | 131,072 | together | moderation | | Llama-3-70b-chat-hf | 8,192 | 8,192 | together | llm | | Llama-3.2-3B-Instruct-Turbo | 131,072 | 131,072 | together | llm | | Meta-Llama-3.1-405B-Instruct-Turbo | 130,815 | 130,815 | together | llm | | scb10x-llama3-typhoon-v1-5x-4f316 | 8,192 | 8,192 | together | llm | | Meta-Llama-3.1-8B-Instruct-Turbo-128K | 131,072 | 131,072 | together | llm | | Meta-Llama-3.1-8B-Instruct-Turbo | 131,072 | 131,072 | together | llm | | Llama-2-13b-chat-hf | 4,096 | 4,096 | together | llm | | Llama-3.2-90B-Vision-Instruct-Turbo | 131,072 | 131,072 | together | llm | | Meta-Llama-3.1-70B-Instruct-Turbo | 131,072 | 131,072 | together | llm | | Meta-Llama-3-8B-Instruct-Turbo | 8,192 | 8,192 | together | llm | | Meta-Llama-3-8B-Instruct-Lite | 8,192 | 8,192 | together | llm | | scb10x-llama3-typhoon-v1-5-8b-instruct | 8,192 | 8,192 | together | llm | | Llama-2-70b-hf | 4,096 | 4,096 | together | llm | | LlamaGuard-2-8b | 8,192 | 8,192 | together | moderation | | Llama-Guard-7b | 4,096 | 4,096 | together | moderation | | Meta-Llama-3-70B-Instruct-Turbo | 8,192 | 8,192 | together | llm | | Meta-Llama-3-70B-Instruct-Lite | 8,192 | 8,192 | together | llm | | Llama-Rank-V1 | 8,192 | 8,192 | together | llm | | Llama-2-7b-chat-hf | 4,096 | 4,096 | together | llm | | Meta-Llama-Guard-3-8B | 8,192 | 8,192 | together | moderation | | llama3-1-8b-instruct-v1 | 128,000 | 128,000 | bedrock | llm | | llama3-1-70b-instruct-v1 | 128,000 | 128,000 | bedrock | llm | | llama3-2-11b-instruct-v1 | 128,000 | 128,000 | bedrock | llm | | llama3-2-90b-instruct-v1 | 128,000 | 128,000 | bedrock | llm | | llama3-2-1b-instruct-v1 | 128,000 | 128,000 | bedrock | llm | | llama3-2-3b-instruct-v1 | 8,192 | 8,092 | bedrock | llm | | llama3-3-70b-instruct-v1 | 8,192 | 8,092 | bedrock | llm | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a model from the **Llama Series**, such as `llama-3-70b-instruct`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "llama-3-70b-instruct", "max_tokens": 150, "messages": [ { "role": "user", "content": "Can you explain how photosynthesis works?" } ] }' ``` --- ### Response Example The expected response structure for the Llama model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "llama-3-70b-instruct", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Photosynthesis is the process by which green plants, algae, and some bacteria convert light energy into chemical energy. Here's how it works:\n\n1. **Light Absorption**: Plants capture light energy using a pigment called chlorophyll, which is found in chloroplasts.\n\n2. **Water and Carbon Dioxide**: They absorb water through their roots and carbon dioxide from the air.\n\n3. **Glucose Production**: The light energy is used to convert water and carbon dioxide into glucose (a sugar) and oxygen. The equation is:\n \n 6CO2 + 6H2O + light energy → C6H12O6 + 6O2\n\nThis process provides energy for the plant and releases oxygen into the atmosphere." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 15, "completion_tokens": 125, "total_tokens": 140, "prompt_characters": 45, "response_characters": 520, "cost": 0.00225, "latency_ms": 3100 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the Llama Series, such as `llama-3-70b-instruct` or others suited to your task. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/features/completions). This example demonstrates how to seamlessly query models from the Llama Series for conversational or instructional tasks. --- ### **Applications and Integrations** - **Conversational AI:** Powering chatbots, virtual assistants, and other dialogue-based systems. Try it with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) or [OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI). - **Code Generation:** Using [CodeLlama](https://ai.meta.com/blog/code-llama-large-language-model-coding/){rel=""nofollow""} variants for programming tasks, code completion, and technical documentation. - **Content Moderation:** Leveraging [LlamaGuard](https://ai.meta.com/research/publications/llama-guard-llm-based-input-output-safeguard-for-human-ai-conversations/){rel=""nofollow""} models for content filtering and safety checks. - **Vision Tasks:** Using vision-enabled models for image understanding and multimodal applications. - **Extended Context Tasks:** Processing long documents with models supporting up to 131K tokens. Learn more in our [Models Guide](https://apipie.ai/docs/features/models). --- ### **Ethical Considerations** The Llama models are powerful tools that should be used responsibly. Users should implement appropriate safeguards and consider potential biases in model outputs. For guidance on responsible AI usage, see Meta's [AI Responsibility Guidelines.](https://ai.meta.com/llama/responsible-use-guide/){rel=""nofollow""} --- ### **Licensing** The Llama Series is available under [Meta's community license agreement](https://github.com/facebookresearch/llama/blob/main/LICENSE){rel=""nofollow""}, which allows for both research and commercial use under specific terms. For detailed licensing information, consult the [official Llama repository](https://github.com/facebookresearch/llama){rel=""nofollow""}, [model cards](https://huggingface.co/meta-llama){rel=""nofollow""}, and respective hosting providers. ::tip Try out the Llama models in APIpie's various \[supported integrations.]\(/docs/Sandbox/Chat) :: # Microsoft AI Models Guide: Phi Series & WizardLM ::div{.docs-image-row} ![Microsoft AI](https://apipie.ai/img/docs/models/Microsoft.png){width="100%"} :: ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Model Comparison and Monitoring** When choosing between Microsoft models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (OpenRouter, Monster API, Together) - Latency metrics and historical performance data - Response time comparisons between different model versions - Specific performance tracking for Phi and WizardLM variants **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different providers and model variants - Usage-based optimization recommendations - Provider-specific pricing tracking **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different capabilities This monitoring system helps users: - Compare costs and pricing across different Microsoft models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose optimal provider based on performance and cost needs ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all Microsoft models. :: ### **Description** The [Microsoft Phi Series](https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/){rel=""nofollow""} represents Microsoft's innovative approach to efficient and powerful language models. Developed by [Microsoft Research](https://www.microsoft.com/en-us/research/){rel=""nofollow""}, these models demonstrate exceptional performance through advanced scaling techniques and architectural improvements. The models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing), offering flexibility in deployment options. #### **Key Features** - **Advanced Architecture:** Models leverage [Microsoft's scaling techniques](https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/){rel=""nofollow""} for optimal performance with smaller parameter counts - **Extended Context Windows:** Support for context lengths from 2K to 128K tokens, with the latest Phi-3 series offering enhanced long-context understanding - **Multi-Provider Availability:** Accessible through platforms like [OpenRouter](https://openrouter.ai/){rel=""nofollow""} and [Monster API](https://monsterapi.ai/){rel=""nofollow""} - **Diverse Applications:** Optimized for chat, instruction-following, and complex reasoning tasks with state-of-the-art performance - **Resource Efficiency:** Industry-leading performance-to-size ratio, making them ideal for cost-effective deployments - **Research-Backed Development:** Built on Microsoft's extensive [research in efficient ML](https://www.microsoft.com/en-us/research/focus-area/ai-and-microsoft-research/){rel=""nofollow""} and model scaling --- ### **Model List in the Microsoft Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::note For detailed performance analysis and benchmarks, see the [Microsoft Phi-2 Technical Report](https://arxiv.org/abs/2311.16716){rel=""nofollow""} and [Phi-2 Research Updates](https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/){rel=""nofollow""}. :: \| **Model Name** | **Max Tokens** | **Response Tokens** | **Providers** | **Subtype** | \| --------------------------------- | -------------- | ------------------- | ------------- | -------------------- | --------------- | \| resnet-50 | - | - | deepinfra | image-classification | \| beit-base-patch16-224-pt22k-ft22k | - | - | deepinfra | image-classification | \| wizardlm-2-8x22b | 65536 | 4096 | openrouter | chatx | text-generation | \| WizardLM-2-7B | 32000 | 32000 | deepinfra | chatx | text-generation | \| WizardLM-2-8x22B | 65536 | 65536 | deepinfra | chatx | text-generation | \| wizardlm-2-7b | 32000 | 4096 | openrouter | text-generation | chatx | \| phi-3-mini-128k-instruct | 128000 | 128000 | openrouter | chatx | \| phi-3-medium-128k-instruct | 128000 | 128000 | openrouter | chatx | \| phi-3.5-mini-128k-instruct | 128000 | 128000 | openrouter | chatx | \| nova-micro-v1 | 128000 | 5120 | bedrock | | \| nova-micro-v1 | 128000 | 5120 | openrouter | chat | \| WizardLM-2-8x22B | 65536 | 65536 | together | chat | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a model from the **Microsoft Series**, such as `phi-3-medium-128k-instruct`. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "phi-3-medium-128k-instruct", "max_tokens": 150, "messages": [ { "role": "user", "content": "What are the key principles of machine learning?" } ] }' ``` --- ### Response Example The expected response structure for the Microsoft model might look like this: ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "phi-3-medium-128k-instruct", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "The key principles of machine learning include:\n\n1. **Data Quality**: High-quality, diverse training data is essential\n\n2. **Feature Selection**: Identifying relevant input variables\n\n3. **Model Selection**: Choosing appropriate algorithms for the task\n\n4. **Training & Validation**: Using separate datasets to ensure generalization\n\n5. **Evaluation**: Measuring performance with appropriate metrics\n\nThese principles form the foundation of effective machine learning systems." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 12, "completion_tokens": 98, "total_tokens": 110, "prompt_characters": 45, "response_characters": 420, "cost": 0.00185, "latency_ms": 2800 }, "system_fingerprint": "fp_123abc456def" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Use any model from the Microsoft Series, such as `phi-3-medium-128k-instruct` or others suited to your task. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response token count (e.g., 150 in this example). - **Messages:** Format your request with a sequence of messages, including user input and system instructions. See [message formatting](https://apipie.ai/docs/features/completions). This example demonstrates how to seamlessly query models from the Microsoft Series for conversational or instructional tasks. --- ### **Applications and Integrations** - **Efficient AI Solutions:** Ideal for applications requiring high performance with resource constraints, leveraging Microsoft's optimized architecture - **Research and Education:** Perfect for academic and research applications, with comprehensive documentation and research papers. - **Extended Context Processing:** Latest models support up to 128K tokens for advanced document analysis and long-form content understanding - **Enterprise Applications:** Production-ready for business applications with Microsoft's enterprise-grade reliability - **Development and Testing:** Excellent for rapid prototyping and development with fast inference times and consistent outputs - **Natural Language Processing:** State-of-the-art performance in text understanding, generation, and analysis tasks --- ### **Ethical Considerations** Microsoft's AI models should be used responsibly with appropriate safeguards and consideration of potential biases. For guidance on responsible AI usage, see Microsoft's [Responsible AI Standards](https://www.microsoft.com/en-us/ai/responsible-ai){rel=""nofollow""}. --- ### **Licensing** The Microsoft Phi Series models are available under specific terms and conditions. For detailed licensing information, consult the [Microsoft Research Open Source](https://opensource.microsoft.com/){rel=""nofollow""} documentation. ::tip Try out the Microsoft models in APIpie's various \[supported integrations.]\(/docs/Sandbox/Chat) :: # Explore Mistral Models & APIs with APIpie Overview ::div{.docs-image-row} ![Mistral AI](https://apipie.ai/img/docs/models/Mistral.png){width="100%"} :: ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Model Comparison and Monitoring** When choosing between Mistral models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (OpenRouter, Together, Deepinfra, Bedrock, EdenAI) - Latency metrics and historical performance data - Response time comparisons between different model versions - Specific performance tracking for Mixtral and Mistral variants **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different providers and model variants - Usage-based optimization recommendations - Provider-specific pricing tracking **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different capabilities (chat, code, embeddings) This monitoring system helps users: - Compare costs and pricing across different Mistral models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose optimal provider based on performance and cost needs ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all Mistral models. :: ### **Description** [Mistral AI](https://mistral.ai/){rel=""nofollow""}, a leading European AI company founded by former [DeepMind](https://deepmind.google/){rel=""nofollow""} and [Meta AI](https://ai.meta.com/){rel=""nofollow""} researchers, develops powerful large language models known for their efficiency and performance. Their models, including the groundbreaking [Mixtral-8x7B](https://mistral.ai/news/mixtral-of-experts/){rel=""nofollow""} (which outperforms GPT-3.5)(one of the best 7B models), consistently achieve top rankings in [open-source model benchmarks](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard){rel=""nofollow""}. The [Mixtral architecture](https://mistral.ai/news/mixtral-of-experts/){rel=""nofollow""} uses a Mixture of Experts approach, allowing it to achieve remarkable performance while maintaining computational efficiency. These models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). #### **Key Features** - **Extended Token Capacity:** Models support context lengths from 4K to 256K tokens, with specialized variants like [Codestral-Mamba](https://huggingface-co.translate.goog/mistralai/Mamba-Codestral-7B-v0.1){rel=""nofollow""} (256K tokens) and [Mistral-Nemo](https://huggingface-co.translate.goog/mistralai/Mistral-Nemo-Instruct-2407){rel=""nofollow""} (128K tokens). - **Multi-Provider Availability:** Accessible through [OpenRouter](https://openrouter.ai/){rel=""nofollow""}, [Together](https://www.together.ai/){rel=""nofollow""}, [Deepinfra](https://deepinfra.com/){rel=""nofollow""}, [Bedrock](https://aws.amazon.com/bedrock/){rel=""nofollow""}, [EdenAI](https://www.edenai.co/){rel=""nofollow""}. - **Diverse Applications:** Optimized for chat, instruction-following, code generation, and embeddings, with specialized models for different tasks. - **Efficiency:** High performance with optimized parameter counts and innovative architectures like Mixture of Experts and Mamba. --- ### **Model List in the Mistral Series** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::note For information about model performance and benchmarks, see the [Mistral AI Documentation](https://docs.mistral.ai/){rel=""nofollow""} and [Model Cards](https://huggingface.co/mistralai){rel=""nofollow""}. :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Providers** | **Subtype** | | --------------------------- | -------------- | ------------------- | ------------------------------- | ----------- | | **Mixtral Models** | | | | | | Mixtral-8x22B-Instruct-v0.1 | 65,536 | 65,536 | Deepinfra, Together, OpenRouter | Chat | | Mixtral-8x7B-Instruct-v0.1 | 32,768 | 32,768 | Deepinfra, Together | Chat | | Mixtral-8x7B-v0.1 | 32,768 | 32,768 | Together | Language | | mixtral-8x7b-instruct | 32,768 | 32,768 | OpenRouter | Chat | | mixtral-8x7b | 32,768 | 32,768 | OpenRouter | Chat | | mixtral-8x7b-instruct-v0 | 4,096 | 4,096 | Bedrock | Chat | | **Mistral Models** | | | | | | Mistral-7B-Instruct-v0.1 | 4,096 | 4,096 | Deepinfra, Together, OpenRouter | Chat | | Mistral-7B-Instruct-v0.2 | 32,768 | 32,768 | Deepinfra, Together, Monster | Chat | | Mistral-7B-Instruct-v0.3 | 32,768 | 4,096 | Together, OpenRouter | Chat | | Mistral-7B-v0.1 | 4,096 | 4,096 | Together | Language | | mistral-7b-instruct | 32,768 | 32,768 | OpenRouter | Chat | | mistral-7b-instruct-v0 | 8,192 | 8,192 | Bedrock | Chat | | mistral-large | 128,000 | 128,000 | OpenRouter | Chat | | mistral-large | 8,192 | 8,192 | Bedrock | Chat | | mistral-medium | 32,000 | 32,000 | OpenRouter | Chat | | mistral-small | 32,000 | 32,000 | OpenRouter | Chat | | mistral-tiny | 32,000 | 32,000 | OpenRouter | Chat | | mistral-small-2402-v1 | 8,192 | 8,192 | Bedrock | Chat | | mistral-nemo | 128,000 | 4,096 | OpenRouter | Chat | | **Specialized Models** | | | | | | codestral-mamba | 256,000 | 256,000 | OpenRouter | Code | | openhermes-2.5-mistral-7b | 4,096 | 4,096 | OpenRouter | Chat | | pixtral-12b | 4,096 | 4,096 | OpenRouter | Chat | | pixtral-large-2411 | 128,000 | 128,000 | OpenRouter | Chat | | ministral-8b | 128,000 | 128,000 | OpenRouter | Chat | | ministral-3b | 128,000 | 128,000 | OpenRouter | Chat | | **Other Variants** | | | | | | 1024\_\_mistral-embed | - | - | EdenAI | Embedding | | large-latest | - | - | EdenAI | Chat | | tiny | - | - | EdenAI | Chat | | small | - | - | EdenAI | Chat | | ministral-3b-latest | - | - | EdenAI | Chat | | ministral-8b-latest | - | - | EdenAI | Chat | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) to interact with a Mistral model: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "mixtral-8x7b-instruct", "max_tokens": 150, "messages": [ { "role": "user", "content": "What are the key differences between traditional and quantum computing?" } ] }' ``` --- ### Response Example ```json { "id": "chatcmpl-12345example12345", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "mixtral-8x7b-instruct", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Here are the key differences between traditional and quantum computing:\n\n1. **Information Processing**\n- Traditional: Uses bits (0 or 1)\n- Quantum: Uses qubits (can be 0, 1, or both simultaneously)\n\n2. **Computation Method**\n- Traditional: Sequential processing\n- Quantum: Parallel processing through superposition\n\n3. **Processing Power**\n- Traditional: Linear scaling\n- Quantum: Exponential scaling for certain tasks\n\n4. **Applications**\n- Traditional: General-purpose computing\n- Quantum: Specialized tasks like cryptography and complex simulations" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 18, "completion_tokens": 132, "total_tokens": 150 } } ``` --- ### API Highlights - **Provider Selection:** Choose from multiple providers or use [automatic routing](https://apipie.ai/docs/features/routing) - **Model Options:** Select from various Mistral models based on your needs - **Token Management:** Set appropriate token limits based on model capabilities - **Message Formatting:** Structure your prompts effectively using [message formatting](https://apipie.ai/docs/features/completions) --- ### **Applications and Integrations** - **Conversational AI:** Build chatbots and virtual assistants. - **Code Generation:** Leverage specialized models like Codestral-Mamba for programming tasks - **Enterprise Solutions:** Deploy models for business applications - **Extended Context Processing:** Handle long documents with models supporting up to 256K tokens - **Research Applications:** Utilize [open-weights models](https://github.com/mistralai){rel=""nofollow""} for academic research and development --- ### **Ethical Considerations** Mistral models should be used responsibly with appropriate safeguards. For guidance on responsible AI usage, see Mistral's [terms](https://mistral.ai/terms/){rel=""nofollow""}. --- ### **Licensing** Mistral models are available under various licensing terms, including the [Apache 2.0 License](https://www.apache.org/licenses/LICENSE-2.0){rel=""nofollow""} for open-source models. For detailed information, consult [Mistral's documentation](https://docs.mistral.ai/){rel=""nofollow""}. ::tip Try out Mistral models in APIpie's [supported integrations](https://apipie.ai/docs/Sandbox/Chat). :: # Flux Image Models Overview ::div{.docs-image-row} ![Flux Models](https://apipie.ai/img/docs/models/FLUX.png){width="50%"} :: ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/images). :: ### **Description** The [Flux image models](https://blackforestlabs.ai/ultra-home/){rel=""nofollow""}, represent a powerful suite of image generation and manipulation models. These models are designed to handle various image creation tasks with exceptional quality and control, offering enterprise-grade reliability through [APIpie's routing system](https://apipie.ai/docs/features/routing). #### **Model Families** **Flux Series:** - **Pro Models:** FLUX.1-pro and FLUX.1.1-pro for highest quality image generation - **Specialized Models:** - FLUX.1-depth for depth-aware image generation - FLUX.1-canny for edge-guided generation - FLUX.1-redux for optimized performance - **Performance Models:** - FLUX.1-schnell series for faster generation - FLUX.1-dev for development and testing #### **Key Features** - **Diverse Generation Capabilities:** Multiple models for different image generation needs - **Specialized Processing:** Models optimized for specific tasks like depth-aware and edge-guided generation - **Performance Options:** From high-quality pro versions to fast schnell variants - **Enterprise Focus:** Built for production-grade image generation applications --- ### **Model List** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::note For detailed information about model capabilities and performance, visit the [Flux Github](https://github.com/black-forest-labs/flux){rel=""nofollow""}. :: | **Model Name** | **Provider** | **Type** | **Description** | **Best For** | | ------------------------- | ------------ | -------- | --------------------------------------------------------------- | ------------------- | | FLUX.1-redux | Together | Image | Optimized general-purpose model | Balanced use | | FLUX.1-dev | Together | Image | Development and testing variant | Testing | | FLUX.1-depth | Together | Image | Depth-aware image generation | 3D-aware images | | FLUX.1-pro | Together | Image | Professional quality generation | High quality | | FLUX.1-schnell | Together | Image | Fast generation model | Quick results | | FLUX.1-schnell-Free | Together | Image | Free tier fast generation | Testing | | FLUX.1-canny | Together | Image | Edge-guided image generation | Control | | FLUX.1.1-pro | Together | Image | Latest professional variant | Best quality | | FLUX-pro | DeepInfra | Image | Flagship model based on Flux latent rectified flow transformers | Professional use | | FLUX-1.1-pro | DeepInfra | Image | Latest state-of-the-art proprietary model | Highest quality | | FLUX-1-dev | DeepInfra | Image | 12B parameter model for text-to-image | Anatomical accuracy | | FLUX-1-schnell | DeepInfra | Image | Fast generation in 1-4 steps | Rapid generation | | FLUX-1-Redux-dev | DeepInfra | Image | Image variation and refinement | Image editing | | Juggernaut-Flux | DeepInfra | Image | Enhanced Flux with sharper details | Detail-rich images | | Juggernaut-Lightning-Flux | DeepInfra | Image | 5x faster high-quality generation | Rapid ideation | --- ### Example API Call Below is an example of how to use the [Image Generation API](https://apipie.ai/docs/features/images) with a Flux model: ```bash curl -L -X POST 'https://apipie.ai/v1/images/generations' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "together", "model": "FLUX.1-pro", "prompt": "A serene landscape with mountains reflected in a calm lake at sunset", "n": 1, "size": "1024x1024" }' ``` --- ### Response Example The expected response structure for a Flux image model: ```json { "created": 1729535643, "data": [ { "url": "https://example.com/generated-image.png", "b64_json": null } ], "usage": { "prompt_tokens": 15, "total_tokens": 15, "cost": 0.008, "latency_ms": 5200 } } ``` --- ### API Highlights - **Model Selection:** - Use pro variants (FLUX.1-pro, FLUX.1.1-pro) for highest quality - Use schnell variants for faster generation - Use specialized variants (depth, canny) for specific needs - **Parameters:** - Customize image size, quality, and generation parameters - Support for various prompting techniques - **Response:** Receive generated images as URLs or base64 encoded data --- ### **Applications and Use Cases** - **Creative Applications:** - Digital art creation - Concept visualization - Design prototyping - Marketing materials - **Professional Use:** - Product visualization - Architectural rendering - Marketing asset creation - Content generation - **Specialized Tasks:** - Depth-aware image generation - Edge-guided generation - Fast prototyping - High-quality final renders --- ### **Ethical Considerations** Flux models should be used responsibly with appropriate content filtering and bias consideration. --- ### **Licensing** Flux models are available through license [available to view on github](https://github.com/black-forest-labs/flux/blob/main/LICENSE){rel=""nofollow""}. ::tip Try out the Flux models in APIpie's various \[supported integrations.]\(/docs/Sandbox/Chat) :: # Leonardo AI Image Models Guide ::div{.docs-image-row} ![Leonardo AI Image Models](https://apipie.ai/img/docs/models/Leonardo_Ai.svg){width="100%"} :: ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/images). :: ### **Description** [Leonardo AI](https://leonardo.ai/){rel=""nofollow""} represents a powerful suite of image generation and manipulation models. These models are designed to handle various image creation tasks with exceptional quality and control, offering enterprise-grade reliability through [APIpie's routing system](https://apipie.ai/docs/features/routing). #### **Model Families** **Leonardo AI Series:** - **Foundation Models:** Phoenix, Leonardo's proprietary foundation model - **Specialized Models:** - Kino XL for cinematic imagery - Leonardo Anime XL for anime-style art - Vision XL for photorealistic outputs - Diffusion XL for versatile image generation - Lightning XL for rapid generation #### **Key Features** - **Diverse Generation Capabilities:** Multiple models for different artistic and photorealistic image needs - **Specialized Processing:** Models optimized for specific styles like cinematic, anime, and photorealistic outputs - **Performance Options:** From high-quality detailed versions to faster generation models - **Enterprise Focus:** Built for production-grade image generation applications - **Customization:** Support for custom fine-tuned models and "Elements" that can be applied to base models --- ### **Model List** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::note For detailed information about model capabilities and performance, visit the [Leonardo AI website](https://leonardo.ai/){rel=""nofollow""}. :: | **Model Name** | **Provider** | **Type** | **Description** | **Best For** | | --------------------- | ------------ | -------- | ------------------------------------- | ---------------- | | Leonardo-Phoenix | EdenAI | Image | Flagship proprietary foundation model | Photorealism | | Leonardo-Lightning-XL | EdenAI | Image | Rapid image generation | Quick results | | Leonardo-Anime-XL | EdenAI | Image | Anime-style art generation | Anime content | | Leonardo-Kino-XL | EdenAI | Image | Cinematic-style imagery | Film-like scenes | | Leonardo-Vision-XL | EdenAI | Image | Photorealistic outputs | Detailed realism | | Leonardo-Diffusion-XL | EdenAI | Image | Versatile image generation | General purpose | --- ### Example API Call Below is an example of how to use the [Image Generation API](https://apipie.ai/docs/features/images) with a Leonardo AI model: ```bash curl -L -X POST 'https://apipie.ai/v1/images/generations' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "edenai", "model": "Leonardo-Phoenix", "prompt": "A serene landscape with mountains reflected in a calm lake at sunset", "n": 1, "size": "1024x1024" }' ``` --- ### Response Example The expected response structure for a Leonardo AI image model: ```json { "created": 1729535643, "data": [ { "url": "https://example.com/generated-image.png", "b64_json": null } ], "usage": { "prompt_tokens": 15, "total_tokens": 15, "cost": 0.017, "latency_ms": 4800 } } ``` --- ### API Highlights - **Model Selection:** - Use Phoenix for highest quality photorealistic imagery - Use Kino XL for cinematic compositions - Use Vision XL for detailed portraits and realistic scenes - Use Anime XL for anime and stylized art - Use Lightning XL for rapid generation - **Parameters:** - Customize image size, quality, and generation parameters - Support for various prompting techniques - **Response:** Receive generated images as URLs or base64 encoded data --- ### **Detailed Model Information** #### **🔥 Phoenix – Leonardo's Proprietary Foundation Model** Phoenix is Leonardo AI's flagship model, developed in-house as a foundational model rather than a fine-tuned variant of existing architectures like Stable Diffusion. It is engineered to deliver high-resolution, photorealistic images with enhanced coherence and detail. #### **🎥 Kino XL – Cinematic Imagery** Kino XL specializes in producing cinematic-style images, excelling in wider aspect ratios and delivering dramatic compositions. It operates effectively without the need for negative prompts, making it user-friendly for creating movie-like scenes. #### **📸 Leonardo Vision XL – Photorealism and Detail** Leonardo Vision XL is tailored for photorealistic outputs, particularly excelling in portraiture and scenes requiring intricate details. It performs best with longer, descriptive prompts, capturing subtle nuances in subjects. #### **🖼️ Leonardo Diffusion XL – Versatile and Efficient** Leonardo Diffusion XL is a versatile model that generates high-quality images even with concise prompts. It's suitable for a broad range of styles and subjects, offering users flexibility in image creation. #### **🖌️ Leonardo Anime XL – Stylized Animation** Leonardo Anime XL focuses on generating anime and manga-style images with consistent character designs, stylized features, and appropriate aesthetic choices that match anime conventions. #### **⚡ Leonardo Lightning XL – Fast Generation** Leonardo Lightning XL provides rapid image generation while maintaining quality, ideal for quick iterations and concept exploration when time is of the essence. --- ### **Applications and Use Cases** - **Creative Applications:** - Digital art creation - Concept visualization - Character design - Storyboarding - **Professional Use:** - Product visualization - Advertising and marketing - Entertainment and media - Game development assets - **Specialized Tasks:** - Photorealistic imagery - Cinematic scene creation - Stylized art generation - Animation and manga-style art --- ### **Ethical Considerations** Leonardo AI models should be used responsibly with appropriate content filtering and bias consideration. Users should be aware of the platform's terms of service and ensure they have appropriate rights for any images used in training custom models. --- ### **Licensing** Leonardo AI models are available through the Leonardo AI platform with specific licensing terms. Refer to [Leonardo AI's terms of service](https://leonardo.ai/terms-of-service){rel=""nofollow""} for detailed information. ::tip Try out the Leonardo AI models in APIpie's various [supported integrations.](https://apipie.ai/docs/Sandbox/Chat) :: # Stable Diffusion Guide: SDXL & Variants ::div{.docs-image-row} ![Stable Diffusion Models](https://apipie.ai/img/docs/models/Stability_AI.png){width="100%"} :: ::info For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/images). :: ### **Model Comparison and Monitoring** When choosing between Stable Diffusion models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (Prodia, Bedrock, DeepInfra, Together) - Latency metrics and historical performance data - Response time comparisons between different model versions - Specific performance tracking for SDXL and specialized variants **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different providers and model variants - Usage-based optimization recommendations - Provider-specific pricing tracking **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different capabilities (base models, SDXL, specialized variants) This monitoring system helps users: - Compare costs and pricing across different Stable Diffusion models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose optimal model based on performance and cost needs ::tip Use the [Models Route](https://apipie.ai/docs/features/models) to access real-time pricing and performance data for all Stable Diffusion models. :: ### **Description** [Stable Diffusion](https://stability.ai/stable-diffusion){rel=""nofollow""}, developed by [Stability AI](https://stability.ai/){rel=""nofollow""}, represents a state-of-the-art collection of image generation models. These models excel at creating high-quality images from text descriptions, offering enterprise-grade reliability through [APIpie's routing system](https://apipie.ai/docs/features/routing). The technology is built on [CompVis's Latent Diffusion](https://github.com/CompVis/latent-diffusion){rel=""nofollow""} architecture and has been trained on [LAION's extensive dataset](https://laion.ai/blog/laion-5b/){rel=""nofollow""}. For community resources and discussions, visit the [Stable Diffusion subreddit](https://www.reddit.com/r/StableDiffusion/){rel=""nofollow""} or join the [official Discord](https://discord.com/invite/stablediffusion){rel=""nofollow""}. The models have been extensively [benchmarked and compared](https://arxiv.org/abs/2302.04222){rel=""nofollow""} in various studies. #### **Technical Specifications** - **Model Architecture:** Based on [Latent Diffusion Models](https://arxiv.org/abs/2112.10752){rel=""nofollow""} - **Training Data:** Trained on [LAION-5B](https://laion.ai/blog/laion-5b/){rel=""nofollow""} dataset - **Model Weights:** Available on [HuggingFace](https://huggingface.co/stabilityai){rel=""nofollow""} - **Compute Requirements:** Varies by model size detailed specifications #### **Performance Metrics** - **FID Score:** Competitive scores across model versions ([benchmark results](https://arxiv.org/abs/2304.06140){rel=""nofollow""}) - **Generation Speed:** From 1-7 seconds depending on hardware and model ([speed comparison](https://huggingface.co/papers/2304.06140){rel=""nofollow""}) - **Resolution Support:** Up to 2048x2048 pixels in SDXL models ([resolution study](https://arxiv.org/abs/2307.01952){rel=""nofollow""}) #### **Model Families** **Stable Diffusion Series:** - **SDXL Models:** Latest generation including animagineXL, dreamshaperXL, and sd\_xl\_base - **SD V1-V2 Models:** Classic models including v1-4, v1-5, and v2-1 - **Specialized Models:** - Realistic Vision series for photorealism - Anime-focused models like dreamlike-anime - Artist-specific models and style variants - Children's story models - Epic realism series - Dreamshaper variants #### **Key Features** - **Diverse Generation Capabilities:** Over 100 models for different artistic styles and use cases - **Provider Options:** Available through multiple providers including Prodia, Bedrock, DeepInfra, and Together - **Performance Options:** From high-quality SDXL to efficient SD base models - **Enterprise Focus:** Built for production-grade image generation applications --- ### **Model List** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** | **Model Name** | **Provider** | **Type** | **Description** | **Best For** | | ------------------------------------ | ------------ | -------- | ------------------------------------ | ----------------- | | 3Guofeng3\_v34 | Prodia | Image | Specialized artistic style | Artistic | | absolutereality\_V16 | Prodia | Image | Photorealistic generation | Realism | | absolutereality\_v181 | Prodia | Image | Enhanced realism | Realism | | amIReal\_V41 | Prodia | Image | Reality simulation | Photorealism | | analog-diffusion-1.0 | Prodia | Image | Analog photo style | Vintage | | anythingv3\_0-pruned | Prodia | Image | General purpose | Versatile | | anything-v4.5-pruned | Prodia | Image | Updated general purpose | Versatile | | anythingV5\_PrtRE | Prodia | Image | Latest anything variant | Versatile | | AOM3A3\_orangemixs | Prodia | Image | Style mixing model | Creative | | blazing\_drive\_v10g | Prodia | Image | Dynamic scene generation | Action scenes | | breakdomain\_I2428 | Prodia | Image | Domain transfer | Style transfer | | breakdomain\_M2150 | Prodia | Image | Enhanced domain transfer | Style transfer | | cetusMix\_Version35 | Prodia | Image | Mixed style generation | Hybrid styles | | childrensStories\_v13D | Prodia | Image | Children's book style | Kids content | | childrensStories\_v1SemiReal | Prodia | Image | Semi-realistic children's style | Kids content | | childrensStories\_v1ToonAnime | Prodia | Image | Animated children's style | Kids animation | | Counterfeit\_v30 | Prodia | Image | Style replication | Style matching | | cuteyukimixAdorable\_midchapter3 | Prodia | Image | Cute anime style | Anime | | cyberrealistic\_v33 | Prodia | Image | Cyberpunk realism | Sci-fi realism | | dalcefo\_v4 | Prodia | Image | Artistic style | Creative | | deliberate\_v2 | Prodia | Image | Controlled generation | Precise output | | deliberate\_v3 | Prodia | Image | Enhanced control | Precise output | | dreamlike-anime-1.0 | Prodia | Image | Anime art generation | Anime | | dreamlike-diffusion-1.0 | Prodia | Image | General purpose | Versatile | | dreamlike-photoreal-2.0 | Prodia | Image | Photorealistic output | Photography | | dreamshaper\_6BakedVae | Prodia | Image | Optimized dreamshaper | Performance | | dreamshaper\_7 | Prodia | Image | Updated dreamshaper | Quality | | dreamshaper\_8 | Prodia | Image | Latest dreamshaper | Best quality | | edgeOfRealism\_eorV20 | Prodia | Image | Realistic edge detection | Detail focus | | EimisAnimeDiffusion\_V1 | Prodia | Image | Specialized anime | Anime | | elldreths-vivid-mix | Prodia | Image | Vivid style mixing | Vibrant art | | epicphotogasm\_xPlusPlus | Prodia | Image | Enhanced photography | Photography | | epicrealism\_naturalSinRC1VAE | Prodia | Image | Natural realism | Nature scenes | | epicrealism\_pureEvolutionV3 | Prodia | Image | Pure realism | Realism | | ICantBelieveItsNotPhotography\_seco | Prodia | Image | Photorealistic output | Photography | | indigoFurryMix\_v75Hybrid | Prodia | Image | Furry art style | Character art | | juggernaut\_aftermath | Prodia | Image | Post-processing focus | Effects | | lofi\_v4 | Prodia | Image | Lo-fi aesthetic | Stylized | | lyriel\_v16 | Prodia | Image | Artistic style | Creative | | majicmixRealistic\_v4 | Prodia | Image | Magic realism | Fantasy realism | | mechamix\_v10 | Prodia | Image | Mechanical style | Mecha | | meinamix\_meinaV9 | Prodia | Image | Style mixing | Creative | | meinamix\_meinaV11 | Prodia | Image | Enhanced mixing | Creative | | neverendingDream\_v122 | Prodia | Image | Dreamlike quality | Surreal | | openjourney\_V4 | Prodia | Image | Open source variant | General use | | pastelMixStylizedAnime\_pruned\_fp16 | Prodia | Image | Pastel anime style | Anime | | portraitplus\_V1.0 | Prodia | Image | Portrait generation | Portraits | | protogenx34 | Prodia | Image | General purpose | Versatile | | Realistic\_Vision\_V1.4 | Prodia | Image | Realistic generation | Realism | | Realistic\_Vision\_V2.0 | Prodia | Image | Enhanced realism | Realism | | Realistic\_Vision\_V4.0 | Prodia | Image | Advanced realism | Realism | | Realistic\_Vision\_V5.0 | Prodia | Image | Latest realism | Realism | | Realistic\_Vision\_V5.1 | Prodia | Image | Current realism | Realism | | redshift\_diffusion-V10 | Prodia | Image | Color enhancement | Color focus | | revAnimated\_v122 | Prodia | Image | Animated style | Animation | | rundiffusionFX25D\_v10 | Prodia | Image | 2.5D effects | Dimensional | | rundiffusionFX\_v10 | Prodia | Image | Special effects | Effects | | sdv1\_4 | Prodia | Image | Base SD v1.4 | General use | | v1-5-pruned-emaonly | Prodia | Image | Optimized v1.5 | Efficiency | | v1-5-inpainting | Prodia | Image | Inpainting specialist | Editing | | shoninsBeautiful\_v10 | Prodia | Image | Beautiful style | Aesthetics | | theallys-mix-ii-churned | Prodia | Image | Style mixture | Creative | | timeless-1.0 | Prodia | Image | Timeless aesthetic | Classic style | | toonyou\_beta6 | Prodia | Image | Cartoon style | Animation | | animagineXLV3\_v30 | Prodia | Image | SDXL anime | Anime | | dreamshaperXL10\_alpha2 | Prodia | Image | SDXL dreamshaper | Quality | | dynavisionXL\_0411 | Prodia | Image | SDXL dynamic | Motion | | juggernautXL\_v45 | Prodia | Image | SDXL effects | Effects | | realismEngineSDXL\_v10 | Prodia | Image | SDXL realism | Realism | | realvisxlV40 | Prodia | Image | SDXL vision | Vision | | sd\_xl\_base\_1.0 | Prodia | Image | SDXL base | Foundation | | sd\_xl\_base\_1.0\_inpainting\_0.1 | Prodia | Image | SDXL inpainting | Editing | | turbovisionXL\_v431 | Prodia | Image | SDXL turbo | Speed | | aniverse\_v30 | Prodia | Image | Anime universe | Anime | | devlishphotorealism\_sdxl15 | Prodia | Image | SDXL photorealism | Realism | | sdxl-turbo | DeepInfra | Image | Fast SDXL with Adversarial Diffusion | Rapid generation | | sd3.5 | DeepInfra | Image | 8B parameter base model | Professional use | | sd3.5-medium | DeepInfra | Image | 2.5B parameter optimized model | Consumer hardware | | stable-diffusion-v1-6 | EdenAI | Image | Base SD v1.6 | General use | | stable-diffusion-xl-1024-v1-0 | EdenAI | Image | SDXL base 1.0 | High quality | | AlbedoBase XL | EdenAI | Image | Leonardo.AI SDXL variant | Advanced art | | SDXL 0.9 | EdenAI | Image | Leonardo.AI SDXL pre-release | Creative work | --- ### Example API Call Below is an example of how to use the [Image Generation API](https://apipie.ai/docs/features/images) with a Stable Diffusion model: ```bash curl -L -X POST 'https://apipie.ai/v1/images/generations' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "prodia", "model": "sd_xl_base_1.0", "prompt": "A serene landscape with mountains reflected in a calm lake at sunset", "n": 1, "size": "1024x1024" }' ``` --- ### Response Example The expected response structure for a Stable Diffusion model: ```json { "created": 1729535643, "data": [ { "url": "https://example.com/generated-image.png", "b64_json": null } ], "usage": { "prompt_tokens": 15, "total_tokens": 15, "cost": 0.008, "latency_ms": 5200 } } ``` --- ### API Highlights and Integration - **Model Selection:** - Use SDXL variants for highest quality - Choose specialized models for specific styles (anime, realistic, artistic) - Select provider based on needs (Prodia, Bedrock, DeepInfra, Together) - **Parameters:** - Customize image size, quality, and generation parameters - Support for advanced prompting techniques - **Response:** Receive generated images as URLs or base64 encoded data --- ### **Applications and Success Stories** - **Creative Applications:** - Digital art creation with tools like [Automatic1111's Web UI](https://github.com/AUTOMATIC1111/stable-diffusion-webui){rel=""nofollow""} - Anime and illustration using specialized models - Character design for games and animation - Concept art for production pipelines - **Professional Use:** - Product visualization for e-commerce - Marketing materials and advertising - Stock photo alternatives - Content creation for social media - Architecture and interior design - **Specialized Tasks:** - Photorealistic imaging - Artistic style transfer and remixing - Character generation for gaming - Scene composition for storyboarding - Fashion design and visualization - **Research and Development:** - [AI art research](https://arxiv.org/abs/2112.10752){rel=""nofollow""} and experimentation - Model fine-tuning and customization - Novel image generation techniques - Computer vision applications --- ### **Ethical Considerations** Stable Diffusion models should be used responsibly with appropriate content filtering and bias consideration. For more information, visit [Stability AI's safety Policy](https://stability.ai/safety){rel=""nofollow""}. --- ### **Licensing and Resources** Stable Diffusion models are available under the [CreativeML Open RAIL-M license](https://huggingface.co/spaces/CompVis/stable-diffusion-license){rel=""nofollow""}. **Additional Resources:** - [Official Documentation](https://platform.stability.ai/docs/getting-started){rel=""nofollow""} - [Model Architecture Paper](https://arxiv.org/abs/2112.10752){rel=""nofollow""} - [HuggingFace Model Hub](https://huggingface.co/stabilityai){rel=""nofollow""} - [GitHub Repository](https://github.com/CompVis/stable-diffusion){rel=""nofollow""} - [Community Wiki](https://wiki.installgentoo.com/wiki/Stable_Diffusion){rel=""nofollow""} ::tip Try out the Stable Diffusion models in APIpie's various [supported integrations.](https://apipie.ai/docs/Sandbox/Chat) :: # OpenAI Models: Explore APIpie's Overview ::div{.docs-image-row} ![OpenAI](https://apipie.ai/img/docs/openai-white.png){width="100%"} :: ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Model Comparison and Monitoring** When choosing between OpenAI models, APIpie provides comprehensive monitoring tools to help make informed decisions: **Performance Monitoring:** - Real-time availability tracking across providers (OpenAI, OpenRouter, EdenAI, DeepInfra) - Latency metrics and historical performance data - Response time comparisons between different model versions - Specific performance tracking for GPT-4, GPT-3.5, and O1 variants **Pricing & Cost Analysis:** - Live pricing updates through our [Models Route](https://apipie.ai/docs/features/models) & [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - Cost comparisons across different providers and model variants - Usage-based optimization recommendations - Provider-specific pricing tracking **Health Metrics:** - Global AI health dashboard for real-time status - Provider reliability tracking - Model uptime statistics - Performance monitoring across different capabilities (chat, vision, audio, voice) This monitoring system helps users: - Compare costs and pricing across different OpenAI models and providers - Track performance metrics to optimize response times - Make data-driven decisions based on availability and reliability - Monitor system health and anticipate potential issues - Choose optimal provider based on performance and cost needs ::tip Use the \[Models Route]\(/docs/features/models) to access real-time pricing and performance data for all OpenAI models. :: ### **Description** [OpenAI's](https://openai.com/){rel=""nofollow""} suite of language models represents the cutting edge in artificial intelligence technology. These models, including the renowned [GPT-4,](https://platform.openai.com/docs/models/#gpt-4o){rel=""nofollow""} [GPT-3.5](https://platform.openai.com/docs/models/#gpt-3-5-turbo){rel=""nofollow""} and [o1 series](https://platform.openai.com/docs/models/#o1){rel=""nofollow""}, deliver exceptional performance across a wide range of tasks. The models are available through various providers integrated with [APIpie's routing system](https://apipie.ai/docs/features/routing). The models come in several specialized variants: - **Chat Models:** Standard chat models optimized for dialogue and instruction-following - **ChatX Models:** Enhanced versions with additional capabilities like function calling and structured output - **O1 Models:** Next-generation models offering superior performance and reliability - **Vision Models:** Capable of understanding and analyzing images alongside text - **Audio Models:** Specialized for audio processing and transcription tasks - **Voice Models:** Text-to-speech models with various voice options - **Code Models:** Optimized for programming and technical documentation tasks The O1 series represents OpenAI's latest advancement in language models, offering: - Improved reasoning and problem-solving capabilities with up to 200K token context - Enhanced consistency in outputs through specialized variants (preview, mini) - Better handling of complex instructions with optimized response generation - Superior performance in specialized tasks with provider-specific optimizations - Flexible deployment options across multiple providers #### **Key Features** - **Extended Context Windows:** Models support context lengths from 4K to 128K tokens, enabling processing of extensive documents and conversations. - **Multi-Provider Availability:** Accessible through [OpenAI](https://openai.com/){rel=""nofollow""}, [OpenRouter](https://openrouter.ai/){rel=""nofollow""}, [EdenAI](https://www.edenai.co/){rel=""nofollow""}, and [DeepInfra](https://deepinfra.com/){rel=""nofollow""}. - **Advanced Capabilities:** - Function calling for structured tool use - JSON mode for reliable structured output - Parallel function calling for efficiency - System message control - Reproducible outputs with seeds - Temperature and top\_p controls - O1 optimizations for enhanced performance - **Multimodal Processing:** Support for text, images, and audio in a single conversation --- ### **Model List** #### **Model List updates dynamically please see [the Models Route](https://apipie.ai/docs/features/models#fetching-models) for the up to date list of models** ::note For information about model performance and benchmarks, see [OpenAI's Model Overview](https://platform.openai.com/docs/models/overview){rel=""nofollow""}. :: | **Model Name** | **Max Tokens** | **Response Tokens** | **Providers** | **Type/Subtype** | | ------------------------------------ | -------------- | ------------------- | -------------------------- | --------------------- | | gpt-4-turbo-preview | 128,000 | 4,096 | OpenAI | LLM/ChatX | | gpt-4-turbo | 128,000 | 4,096 | OpenAI, EdenAI | LLM/ChatX | | gpt-4-turbo-2024-04-09 | 4,096 | 4,096 | OpenAI, EdenAI | LLM/Chat | | gpt-4-1106-preview | 4,096 | 4,096 | OpenAI, EdenAI | LLM/Chat | | gpt-4 | 8,191 | 4,096 | OpenAI, OpenRouter, EdenAI | LLM/Chat | | gpt-4-0613 | 8,192 | 8,192 | OpenAI | LLM/ChatX | | gpt-4-0314 | 8,191 | 4,096 | EdenAI | LLM/Chat | | gpt-4-32k-0314 | 32,767 | 4,096 | EdenAI | LLM/Chat | | gpt-4o | 128,000 | 16,384 | OpenAI, OpenRouter, EdenAI | LLM/Chat,ChatX,Code | | gpt-4o-mini | 128,000 | 16,384 | OpenAI, EdenAI | LLM/Chat,ChatX | | gpt-4o-mini-2024-07-18 | 16,384 | 16,384 | OpenAI | LLM/Chat,ChatX | | gpt-4o-2024-05-13 | 4,096 | 4,096 | OpenAI, EdenAI | LLM/Chat,ChatX | | gpt-4o-2024-08-06 | 16,384 | 16,384 | OpenAI | LLM/Chat,ChatX | | gpt-4o-2024-11-20 | 16,384 | 16,384 | OpenAI, OpenRouter | LLM/Chat,ChatX | | o1-preview | 128,000 | 32,768 | OpenAI | LLM/ChatX | | o1-mini | 128,000 | 32,768 | EdenAI | LLM/Chat | | o1 | 200,000 | 100,000 | OpenRouter | LLM | | gpt-4-vision-preview | 128,000 | 4,096 | OpenAI, OpenRouter | Vision/Multimodal | | gpt-4-1106-vision-preview | 128,000 | 4,096 | OpenAI | Vision/ChatX | | gpt-4o-audio-preview | 16,384 | 16,384 | OpenAI | LLM | | gpt-4o-mini-audio-preview | 128,000 | 16,384 | OpenAI | LLM | | gpt-4o-audio-preview-2024-12-17 | 16,384 | 16,384 | OpenAI | LLM | | gpt-4o-mini-audio-preview-2024-12-17 | 128,000 | 16,384 | OpenAI | LLM | | gpt-3.5-turbo | 16,384 | 4,096 | OpenAI, EdenAI | LLM/Chat | | gpt-3.5-turbo-16k | 16,385 | 4,096 | EdenAI | LLM/Chat | | gpt-3.5-turbo-0125 | 16,385 | 4,096 | OpenAI, EdenAI | LLM/Chat | | gpt-3.5-turbo-1106 | 16,385 | 4,096 | OpenAI, EdenAI | LLM/Chat | | gpt-3.5-turbo-instruct | 4,095 | 4,096 | OpenRouter | LLM/Chat | | tts-1-hd shimmer | - | - | OpenAI | Voice/TTS | | tts-1-hd alloy | - | - | OpenAI | Voice/TTS | | tts-1-hd echo | - | - | OpenAI | Voice/TTS | | tts-1-hd fable | - | - | OpenAI | Voice/TTS | | tts-1-hd onyx | - | - | OpenAI | Voice/TTS | | tts-1-hd nova | - | - | OpenAI | Voice/TTS | | tts-1-1106 alloy | - | - | OpenAI | Voice/TTS | | tts-1-1106 echo | - | - | OpenAI | Voice/TTS | | tts-1-1106 fable | - | - | OpenAI | Voice/TTS | | tts-1-1106 onyx | - | - | OpenAI | Voice/TTS | | tts-1-1106 nova | - | - | OpenAI | Voice/TTS | | clip-vit-large-patch14-336 | - | - | DeepInfra | Vision/Classification | | clip-vit-large-patch14 | - | - | DeepInfra | Vision/Classification | | clip-vit-base-patch32 | - | - | DeepInfra | Vision/Classification | | whisper-large | - | - | DeepInfra | Voice/ASR | | whisper-medium | - | - | DeepInfra | Voice/ASR | | whisper-medium.en | - | - | DeepInfra | Voice/ASR | | whisper-timestamped-medium | - | - | DeepInfra | Voice/ASR | | whisper-timestamped-medium.en | - | - | DeepInfra | Voice/ASR | | whisper-base | - | - | DeepInfra | Voice/ASR | | whisper-base.en | - | - | DeepInfra | Voice/ASR | | whisper-small | - | - | DeepInfra | Voice/ASR | | whisper-small.en | - | - | DeepInfra | Voice/ASR | | whisper-tiny | - | - | DeepInfra | Voice/ASR | | whisper-tiny.en | - | - | DeepInfra | Voice/ASR | --- ### Example API Call Below is an example of how to use the [Chat Completions API](https://apipie.ai/docs/features/completions) with OpenAI's GPT-4 model: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openai", "model": "gpt-4-turbo-preview", "max_tokens": 150, "messages": [ { "role": "user", "content": "What are the key differences between renewable and non-renewable energy sources?" } ] }' ``` --- ### Response Example The expected response structure from an OpenAI model: ```json { "id": "chatcmpl-abc123example456", "object": "chat.completion", "created": 1709234567, "provider": "openai", "model": "gpt-4-turbo-preview", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Renewable and non-renewable energy sources differ in several key ways:\n\n1. Replenishment:\n- Renewable: Naturally replenished within a human lifetime (solar, wind, hydro)\n- Non-renewable: Take millions of years to form (fossil fuels)\n\n2. Environmental Impact:\n- Renewable: Generally lower emissions and environmental impact\n- Non-renewable: Higher carbon emissions and environmental degradation\n\n3. Availability:\n- Renewable: Unlimited but dependent on natural conditions\n- Non-renewable: Limited and depleting resources" }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 20, "completion_tokens": 110, "total_tokens": 130, "prompt_characters": 82, "response_characters": 380, "cost": 0.0035, "latency_ms": 2800 }, "system_fingerprint": "fp_abc123def456" } ``` --- ### API Highlights - **Provider:** Specify the provider or leave blank for [automatic selection](https://apipie.ai/docs/features/routing). - **Model:** Choose from OpenAI's range of models based on your needs. See [Models Guide](https://apipie.ai/docs/features/models#fetching-models). - **Max Tokens:** Set the maximum response length (varies by model). - **Messages:** Structure your conversations with role-based messages. See [message formatting](https://apipie.ai/docs/features/completions). --- ### **Applications and Integrations** - **Conversational AI:** Power chatbots and virtual assistants with [GPT-4 and GPT-3.5](https://platform.openai.com/docs/guides/chat){rel=""nofollow""}. Try with [LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat). - **High-Performance Tasks:** Utilize O1 models for applications requiring superior reliability and consistent outputs. - **Vision Tasks:** Process and analyze images with [GPT-4 Vision](https://platform.openai.com/docs/guides/vision){rel=""nofollow""} models. - **Audio Processing:** Handle audio tasks with specialized audio preview models. - **Extended Context:** Process long documents with models supporting up to 128K tokens. See our [Models Guide](https://apipie.ai/docs/features/models). - **Code Generation:** Leverage models for programming tasks and technical documentation. --- ### **Ethical Considerations** OpenAI models are powerful tools that require responsible use. Users should implement appropriate safeguards and consider potential biases. For guidance, refer to OpenAI's [Usage Policies](https://openai.com/policies/usage-policies){rel=""nofollow""} and [Safety Best Practices](https://platform.openai.com/docs/guides/safety-best-practices){rel=""nofollow""}. --- ### **Licensing** OpenAI's models are available under commercial terms through their [API Service Terms](https://openai.com/policies/terms-of-use){rel=""nofollow""}. For detailed licensing information and usage guidelines, consult the [OpenAI Platform Documentation](https://platform.openai.com/docs){rel=""nofollow""} and respective hosting providers. ::tip Try out OpenAI models in APIpie's various [supported integrations](https://apipie.ai/docs/Sandbox/Chat). :: # ElevenLabs Voice AI Guide: TTS and Voice Synthesis ::div{.docs-image-row} ![ElevenLabs Models](https://apipie.ai/img/docs/models/ElevenLabs.png){width="100%"} :: ::info For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ### **Description** [ElevenLabs](https://elevenlabs.io){rel=""nofollow""} is a leading provider of state-of-the-art text-to-speech technology. Their advanced AI models offer natural-sounding voice synthesis with unprecedented quality and control. The platform supports multiple languages and can generate highly realistic speech with various voices, accents, and emotional tones. Learn more about ElevenLabs: - [Official Documentation](https://docs.elevenlabs.io){rel=""nofollow""} - [Voice Library](https://elevenlabs.io/voice-library){rel=""nofollow""} - [Speech Synthesis Blog](https://elevenlabs.io/blog){rel=""nofollow""} - [Discord Community](https://discord.gg/elevenlabs){rel=""nofollow""} - [GitHub Resources](https://github.com/elevenlabs){rel=""nofollow""} ### **Technology Overview** ElevenLabs utilizes cutting-edge deep learning techniques to create natural-sounding synthetic voices. Their technology incorporates: - Neural voice cloning - Multilingual speech synthesis - Real-time voice generation - Emotion and style control - High-fidelity audio output ### **Available Models** | Model | Max Tokens | Provider | Type | Description | | ------------------------ | ---------- | ---------- | ---- | ----------------------------------------------- | | eleven\_multilingual\_v2 | 5000 | elevenlabs | tts | Latest multilingual model with enhanced quality | | eleven\_multilingual\_v1 | 5000 | elevenlabs | tts | First generation multilingual model | | eleven\_monolingual\_v1 | 5000 | elevenlabs | tts | English-optimized model | | eleven\_turbo\_v2 | 5000 | elevenlabs | tts | Fast generation model | | eleven\_turbo\_v2\_5 | 5000 | elevenlabs | tts | Enhanced turbo model | | eleven\_flash\_v2 | 5000 | elevenlabs | tts | Ultra-fast generation | | eleven\_flash\_v2\_5 | 5000 | elevenlabs | tts | Latest ultra-fast model | ### **Available Voices** ElevenLabs provides a diverse set of pre-made voices with different characteristics: #### Professional Narration - **Rachel**: Young female, American accent, calm tone - ideal for narration - **Drew**: Middle-aged male, American accent - perfect for news reading - **Antoni**: Young male, American accent - well-rounded narrator - **Thomas**: Young male, American accent - calm meditation voice - **Bill**: Older male, American accent - trustworthy narration #### Character Voices - **Clyde**: Middle-aged male, American accent - war veteran character - **Dave**: Young male, British-Essex accent - conversational gaming voice - **Fin**: Older male, Irish accent - sailor character - **Glinda**: Middle-aged female, American accent - witch character - **Charlotte**: Young female, Swedish accent - seductive character #### News & Media - **Paul**: Middle-aged male, American accent - ground reporter - **Sarah**: Young female, American accent - soft news voice - **Daniel**: Middle-aged male, British accent - authoritative news - **Alice**: Middle-aged female, British accent - confident news - **Joseph**: Middle-aged male, British accent - field reporter #### Social Media & Casual - **Laura**: Young female, American accent - upbeat social media - **Will**: Young female, American accent - friendly social - **Jessica**: Young female, American accent - expressive conversational - **Eric**: Middle-aged male, American accent - friendly conversational - **Chris**: Middle-aged male, American accent - casual style #### Special Purpose - **Ethan**: Young male, American accent - ASMR/whisper - **Nicole**: Young female, American accent - audiobook whisper - **Dorothy**: Young female, British accent - children's stories - **Michael**: Older male, American accent - audiobook narration - **Grace**: Young female, American Southern accent - gentle audiobook #### Multilingual - **Giovanni**: Young male, Italian-English accent - foreign audiobook - **Mimi**: Young female, Swedish-English accent - animation - **Charlie**: Middle-aged male, Australian accent - conversational - **James**: Older male, Australian accent - news - **George**: Middle-aged male, British accent - warm narration ### Example API Call ```bash curl -X POST 'https://apipie.ai/v1/audio/speech' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer YOUR_API_KEY' \ --data-raw '{ "model": "eleven_multilingual_v2", "voice": "Rachel", "input": "Hello! This is a test of the ElevenLabs text to speech API.", "voice_settings": { "stability": 0.5, "similarity_boost": 0.75 } }' ``` ### Response Example ```json { "created": 1729535643, "audio": { "content_type": "audio/mpeg", "url": "https://example.com/generated-audio.mp3" }, "usage": { "text_characters": 57, "cost": 0.004275, "latency_ms": 1200 } } ``` ### **Voice Settings and Controls** You can customize the voice output using these parameters: - **stability** (0-1): Controls voice stability. Higher values make the voice more consistent but less expressive - **similarity\_boost** (0-1): Enhances similarity to the original voice. Higher values make it sound more like the reference - **style** (0-1): Adjusts speaking style intensity - **use\_speaker\_boost** (boolean): Enhances speaker clarity ### **Integration and Use Cases** ElevenLabs' text-to-speech technology can be integrated into various applications: 1. **Content Creation** - Audiobook production - Podcast generation - Video narration - E-learning content 2. **Entertainment** - Game character voices - Animation dubbing - Interactive storytelling - Voice-enabled NPCs 3. **Business Applications** - IVR systems - Virtual assistants - Customer service - Corporate training 4. **Accessibility** - Screen readers - Text-to-speech for visually impaired - Language learning tools - Reading assistance ### **Best Practices and Optimization** 1. **Model Selection**: - Use multilingual\_v2 for highest quality across languages - Use turbo or flash models for faster generation - Use monolingual for English-only applications 2. **Voice Selection**: - Choose voices based on use case (narration, characters, news, etc.) - Consider accent and age appropriate for your content - Test multiple voices to find the best fit 3. **Text Preparation**: - Use punctuation to control pacing - Break long text into natural segments - Include phonetic spelling for unusual words ### **Performance and Limitations** - Maximum text length varies by subscription - Some voices may have accent or language restrictions - Generation time varies by model and text length ### **Security and Ethics** ElevenLabs maintains strict guidelines for voice usage: - Voice cloning requires explicit consent - Built-in content filtering - Usage monitoring and abuse prevention - Secure API access and authentication ### **Resources and Support** Get help and learn more: - [ElevenLabs Status Page](https://status.elevenlabs.io){rel=""nofollow""} - [API Reference](https://api.elevenlabs.io/docs){rel=""nofollow""} - [Speech Synthesis Guide](https://docs.elevenlabs.io/speech-synthesis){rel=""nofollow""} - [Voice Design Guide](https://docs.elevenlabs.io/voice-design){rel=""nofollow""} ::tip Try out ElevenLabs voices in APIpie's [supported integrations](https://apipie.ai/docs/Sandbox/Chat) :: # User Agreement - Terms of Service *Effective on April 1, 2022 | Last updated on November 14, 2025* ## Terms and Conditions These terms and conditions ("Terms", "Agreement") are an agreement between **Neuronic AI Inc. d/b/a APIpie AI** ("Neuronic AI Inc.", "us", "we" or "our") and you ("User", "you" or "your"). This Agreement sets forth the general terms and conditions of your use of the APIpie platform and any of its products or services (collectively, "**Services**"). This includes, but is not limited to, our cutting-edge AI model aggregation, orchestration, RAG Tuning, and data enhancement services. --- ### 1. Accounts and Membership You must be at least 13 years of age to use the Services. By using the Services and by agreeing to this Agreement, you warrant and represent that you are at least 13 years of age. If you create an account, you are responsible for maintaining the security of your account and API keys and for all activities that occur under the account. You are also responsible for any actions taken in connection with the account. Providing false contact information of any kind may result in the termination of your account. You must notify us immediately of any unauthorized uses of your account or any other breaches of security. We will not be liable for any acts or omissions by you, including any damages incurred as a result of such acts or omissions. We may suspend, disable, or delete your account (or any part thereof) if we determine that you have violated any provision of this Agreement or that your conduct or content would tend to damage our reputation and goodwill. If we delete your account for the aforementioned reasons, you may not re-register for our Services. We may also block your email address and Internet protocol address to prevent further registration. --- ### 2. Billing and Payments You shall pay all fees or charges to your account in accordance with the fees, charges, and billing terms in effect at the time a fee or charge is due and payable. You acknowledge that charges may vary based on the model selected via **Preferred Routing** or **AI Model Pooling**, and based on token consumption tracked accurately by our **Usage** feature. We reserve the right to change service pricing and offerings at any time. We also reserve the right to refuse any order you place with us. In our sole discretion, we may limit or cancel quantities purchased per person, per household, or per order. These restrictions may include orders placed by or under the same customer account, the same credit card, or orders that use the same billing and/or shipping address. If we make a change to or cancel an order, we will attempt to notify you by contacting the email address and/or billing address/phone number provided at the time the order was made. --- ### 3. Accuracy of Information Occasionally, there may be information on the Services that contains typographical errors, inaccuracies, or omissions that relate to product descriptions, pricing, availability, promotions, and offers. We reserve the right to correct any errors, inaccuracies, or omissions and to change or update information or cancel orders if any information on the Services is inaccurate at any time without prior notice (including after you have submitted your order). We undertake no obligation to update, amend, or clarify information on the Services, including pricing information, except as required by law. The absence of a specified update or refresh date applied on the Services should not be taken as an indication that all information on the Services has been modified or updated. --- ### 4. User Content and Data Backups **User Content** includes the inputs (prompts, documents for **RAG Tuning**, files) and outputs generated via the Services. - **Responsibility for Content:** You are solely responsible for the content you submit to the Services, including ensuring you have the necessary rights to use such data. - **Data Retention:** While we implement secure storage for persistent data like **RAGTune Document Collections** and long-term memory vectors (which you control via API), we are not responsible for Content residing on the Services. - **Backups:** In no event shall we be held liable for any loss of any Content. It is your sole responsibility to maintain appropriate backups of your Content. On some occasions and in certain circumstances, with no obligation, we may be able to restore some or all of your data that has been deleted as of a certain date and time when we may have backed up data for our own purposes. However, we make no guarantee that the data you need will be available. --- ### 5. AI Behavior, Data Usage, and Sub-processors In the course of providing the Services, we utilize various AI technologies and Sub-processors to execute features like **Completions**, **Integrity**, and **Image Generation**. - **Accuracy and Reliability:** You understand and accept that AI models, even those enhanced by features like **Integrity** (which minimizes hallucinations), have a margin of error and may not always provide accurate or desired results. You acknowledge that the AI systems are capable of independent operation and we bear no responsibility for the actions, statements, or behaviors of these AI systems, including any costs incurred by the AI's activities on your behalf. - **Data Processing:** Any data provided by you (User Content) is processed according to our **Privacy Policy** and our **Data Processing Addendum (DPA)**, located at [Data Processing Addendum (DPA)](https://apipie.ai/dpa). - **Third-Party Sharing:** You acknowledge that fulfilling API requests requires transmitting User Content to external LLM Providers and other Sub-processors (e.g., vector database providers like Pinecone, if enabled). **By using the Services, you explicitly consent to the transfer of your User Content to these Sub-processors as necessary to provide the features you utilize.** - **Anonymization:** If you utilize our **Anonymization Services** (e.g., using `anon: true` in your prompt), we will employ best-effort processing to redact or anonymize your data prior to transmission to an external LLM, but we cannot guarantee absolute and complete anonymization. --- ### 6. Advertisements During the use of the Services, you may correspond with advertisers or sponsors showing their goods or services through the Services. Any such activity, and any terms, conditions, warranties, or representations associated with such activity, are solely between you and the applicable third-party. We shall have no liability, obligation, or responsibility for any such correspondence, purchase, or promotion between you and any such third-party. --- ### 7. Links to Other Websites Although the Services may link to other websites, we are not implying any approval, association, sponsorship, endorsement, or affiliation with any linked website, unless specifically stated herein. We are not responsible for examining or evaluating the offerings of any businesses or individuals or the content of their websites. We do not assume any responsibility or liability for the actions, products, services, or content of any other third-parties. You should review the legal statements and other conditions of use of any website which you access through a link from the Services. Your linking to any other off-site websites is at your own risk. --- ### 8. Prohibited Uses In addition to other terms set forth in this Agreement, you are prohibited from using the Services or its Content: (a) for any unlawful purpose; (b) to solicit others to perform or participate in any unlawful acts; (c) to violate any international, federal, provincial or state regulations, rules, laws, or local ordinances; (d) to infringe upon or violate our intellectual property rights or the intellectual property rights of others; (e) to harass, abuse, insult, harm, defame, slander, disparage, intimidate, or discriminate based on gender, sexual orientation, religion, ethnicity, race, age, national origin, or disability; (f) to submit false or misleading information; (g) to upload or transmit viruses or any other type of malicious code that will or may be used in any way that will affect the functionality or operation of the Service or of any related website, other websites, or the Internet; (h) to collect or track the personal information of others (except as necessary and compliant with applicable law for your own application's users); (i) to spam, phish, pharm, pretext, spider, crawl, or scrape; (j) for any obscene or immoral purpose; or (k) to interfere with or circumvent the security features of the Service or any related website, other websites, or the Internet. We reserve the right to terminate your use of the Service or any related website for violating any of the prohibited uses. --- ### 9. Intellectual Property Rights This Agreement does not transfer to you any intellectual property owned by Neuronic AI Inc. or third-parties. All rights, titles, and interests in and to such intellectual property will remain solely with Neuronic AI Inc. or the respective licensors. All trademarks, service marks, graphics, and logos used in connection with our Services are trademarks or registered trademarks of Neuronic AI Inc. or their respective licensors. Other trademarks, service marks, graphics, and logos used in connection with our Services may be the trademarks of other third-parties. Your use of our Services grants you no right or license to reproduce or otherwise use any Neuronic AI Inc. or third-party trademarks. --- ### 10. Disclaimer of Warranty You agree that your use of our Services is solely at your own risk. We disclaim all warranties of any kind, whether express or implied, including but not limited to the implied warranties of merchantability, fitness for a particular purpose, and non-infringement. We make no warranty that the Services will meet your requirements or be uninterrupted, timely, secure, or error-free. We also make no warranty as to the results that may be obtained from the use of the Service or the accuracy or reliability of any information obtained through the Service. You understand and agree that any material and/or data downloaded or otherwise obtained through the use of the Service is done at your own discretion and risk, and you will be solely responsible for any damage to your computer system or loss of data that results from the download of such material and/or data. We make no warranty regarding any goods or services purchased or obtained through the Service or any transactions entered into through the Service. No advice or information, whether oral or written, obtained by you from us or through the Service shall create any warranty not expressly made herein. --- ### 11. Changes and Amendments We reserve the right to modify this Agreement or its policies relating to the Services at any time, effective upon posting of an updated version of this Agreement on the Website. When we do, we will revise the updated date at the top of this page. Continued use of the Services after any such changes shall constitute your consent to such changes. --- ### 12. Indemnification You agree to indemnify and hold Neuronic AI and its affiliates, directors, officers, employees, and agents harmless from and against any liabilities, losses, damages or costs, including reasonable attorneys' fees, incurred in connection with or arising from any third-party allegations, claims, actions, disputes, or demands asserted against any of them as a result of or relating to your Content, your use of the Services or any willful misconduct on your part. --- ### 13. Limitation of Liability To the fullest extent permitted by applicable law, in no event will Neuronic AI, its affiliates, officers, directors, employees, agents, suppliers or licensors be liable to any person for any indirect, incidental, special, punitive, cover or consequential damages (including, without limitation, damages for lost profits, revenue, sales, goodwill, use of content, impact on business, business interruption, loss of anticipated savings, loss of business opportunity) however caused, under any theory of liability, including, without limitation, contract, tort, warranty, breach of statutory duty, negligence or otherwise, even if Neuronic AI has been advised as to the possibility of such damages or could have foreseen such damages. --- ### 14. Acceptance of These Terms You acknowledge that you have read this Agreement and agree to all its terms and conditions. By using the Services, you agree to be bound by this Agreement. If you do not agree to abide by the terms of this Agreement, you are not authorized to use or access the Services. # API Agreement: Key Terms & Usage Guidelines *Last updated on November 14, 2025* ### 1. Acceptance The APIpie.ai API and its related features (as described below) are provided by **Neuronic AI Inc. d/b/a APIpie AI** (“Neuronic AI”, “Company”, “we”, “us” or “our”). By using our APIpie.ai API, you must fully accept all the terms and conditions within this APIpie.ai API User Agreement (“API Agreement”). By subscribing to any of our API services, you accept the terms of, and agree to be fully bound by, this API Agreement. In addition, by subscribing to, accessing, or otherwise using our APIpie.ai API, you agree to be legally bound by all provisions of our **Privacy Policy** and **User Agreement**. If you are using our APIpie.ai API or entering into this API Agreement on behalf of a corporation, partnership, or other legal entity, or on behalf of any other third party, you represent and warrant that you are duly authorized to represent and bind such entity or third party to this API Agreement. In this context, all references to “you” or “your” shall also refer to such entity or other third party. Should you lack the aforementioned authority or if you (or the entity that you represent) do not agree with the terms within this API Agreement, our Privacy Policy and/or User Agreement, you are not permitted to use the APIpie.ai API. As a user of the APIpie.ai API, you acknowledge and accept that this API Agreement, our Privacy Policy, and User Agreement constitutes a legally binding contract between you and Neuronic AI. In the event of any conflict between this API Agreement and our Privacy Policy or User Agreement, if any, and as amended from time to time, the terms of this API Agreement shall control. --- ### 2. Definitions 2.1 “Application” or “App” means any software or mobile application, website, product, or service developed, created, or offered using our APIpie.ai API. 2.2 “API Documentation” means the documentation, data, and information that Neuronic AI provides regarding the use of our APIpie.ai API through our APIpie.ai Site. 2.3 “API Site” means our API development site found at {rel=""nofollow""} 2.4 “APIpie.ai API” means the Application Programming Interface (“API”) made available to the public by our Company, including all features such as **Completions**, **Models Route**, **RAG Tuning**, **Integrity**, **AI Model Pooling**, **Preferred Routing**, **Tools Support**, **Image Generation**, **Voice Synthesis**, and **Usage**, as well as the related API Documentation. 2.5 “Neuronic AI Brand” refers to the “Neuronic AI” brand name and other brandings, names, logos, slogans, services marks, trademarks, and trade names belonging to Neuronic AI, whether or not registered, and existing anywhere worldwide. 2.6 “Data” means any data and content uploaded, posted, transmitted, or otherwise made available by Neuronic AI on any of its websites or other associated sites/platforms/channels. 2.7 “Service(s)” refers to Neuronic AI’s products and services, including but not limited to the APIpie.ai API, all of its websites (the “Site”), and all software, applications, mobile applications, text, images, messaging channels, message services, updates, data, reports, content, newsletters, databases, forums, articles, guides, reports, and other information made available by or on behalf of Neuronic AI through any of the foregoing. The “Service” does not include information, software application, mobile application, platform, website, or service provided by you or a third party (including Applications), whether or not Neuronic AI designates them as “official integrations”. 2.8 “User Content” refers to the data, inputs (prompts, documents for RAG, files), and outputs generated by you or your users through the APIpie.ai API. 2.9 “Users” refers to any and all users of our APIpie.ai API, which includes you. --- ### 3. Grant of License Pursuant to the provisions of this API Agreement, Neuronic AI hereby grants you a limited, non-exclusive, non-assignable, non-transferable license to use the APIpie.ai API to develop, test, and support any software application, mobile application, website, platform, service, or product, as well as to integrate or incorporate the APIpie.ai API with your Application. This license granted to you is subject to the limitations set forth herein. You specifically agree that violation of Section 4 below will result in the automatic termination of the license granted to you herein. --- ### 4. Scope of Use 4.1 Your use of APIpie.ai's API is limited to the following: 4.1.1 As part of our commitment to continual development and improvement, changes may be made to the APIpie.ai API or this API User Agreement at any time and without prior notice. This includes modifications to models available via the **Models Route**, changes in routing logic via **Preferred Routing** or **Intelligent Model Management (IMM)**, or updates to the **Integrity** system. You agree that it's your responsibility to regularly check our website for updates. Keep in mind that some parts of our API may not be fully documented. Any reliance on a specific function, behavior, or aspect of our API should not be done without considering the potential for change. Neuronic AI shall not be held liable for any negative outcomes arising from such reliance, as detailed in Section 11 below. 4.1.2 Our property (including but not limited to our APIpie.ai API) should not be used in breach of any laws or regulations, or to infringe on the rights of any person or entity. This includes usage that infringes on a third party's intellectual property rights, privacy rights, or that is inconsistent with this API Agreement, our Privacy Policy, User Agreement, or any other agreements with Neuronic AI. 4.1.3 The APIpie.ai API should not be used to access or use any information beyond what is permitted under this API Agreement, or in a manner that disrupts, interferes with, or degrades the performance of our services, circumvents our security measures, or tests the vulnerability of our systems and networks. 4.1.4 You agree not to introduce any damaging or intrusive computer programs, including but not limited to worms, trojans, viruses, hacks, or other programs that may damage, interfere with, or expropriate any system or data. 4.1.5 You agree not to reverse engineer or derive source codes, trade secrets, or know-how from any of our property or to attempt any of the foregoing. 4.1.6 Neuronic AI may offer products or services that are similar to your offerings in the future, and you agree that we are fully entitled to do so without any restrictions or notice. 4.1.7 You are allowed to charge for your services and products that incorporate our API. However, you are not permitted to sell, rent, lease, sublicense, redistribute, or syndicate access to the APIpie.ai API or any part thereof. 4.1.8 You may place advertisements on and around your products, services, website, platforms, mobile applications, and software applications ("your Products") that incorporate or integrate our API. However, you are not permitted to: Place any advertisements on any of our Property Make any public statements or representations in any mode, media, or channel regarding Neuronic AI or any of our Products (including the APIpie.ai API) without our prior written consent. 4.2 The rate limit for the APIpie.ai API is based on user subscriptions and restricted by API keys. You agree not to exceed or circumvent these limitations or use the API in a way that constitutes excessive or abusive usage. Usage is monitored transparently via the **Usage** feature. 4.3 Except as permitted within this API Agreement, you will not use the Neuronic AI brand in a way that may suggest that your Products are endorsed or sponsored by Neuronic AI without an additional agreement. --- ### 5. Data Handling, Caching, and Sub-processors 5.1 **User Content and Data Processing:** All User Content is handled in accordance with the **Privacy Policy** and the **Data Processing Addendum (DPA)**, located at [Data Processing Addendum (DPA)](https://apipie.ai/dpa). By using the API, you consent to the necessary transfer of User Content to third-party LLM providers and Sub-processors (e.g., Pinecone for **RAG Tuning**) as detailed in the DPA for the sole purpose of providing the Services. 5.2 **Data Anonymization:** If you utilize our pseudo-anonymization or parsing features (e.g., using `anon: true`), you acknowledge that we employ third-party tools (like Presidio and Tika) to attempt redaction, but Neuronic AI cannot guarantee 100% complete and absolute anonymization of all personal data. 5.3 **Data Storage (Non-User Content):** If you must cache or store non-User Content related to the API (Including but not limited to model availability, Cost, Latency): - Refresh the cache at least every 24 hours - Apply strong encryption and other security measures to stored data - Promptly delete all user data collected via the APIpie.ai API at a user's request, unless required for legitimate legal or business purposes. - Promptly and permanently delete all data and other information stored via your use of the APIpie.ai API upon termination of your access to the API. :br 5.4 **Restriction on Duplication:** You are not allowed to duplicate, reproduce, copy, store, derive from, or translate any data, API documentation, or any information expressed by the data, except as expressly permitted under this API Agreement. --- ### 6. User Agreement for Your Products 6.1 If your products are offered for use to others outside of your entity, you must have a binding user agreement and privacy policy in place that: - Identifies the APIpie.ai API as being the property of Neuronic AI - Ensures that your users abide by this API Agreement when using the APIpie.ai API - Excludes and disclaims all liability of Neuronic AI for all usage of the API - Assumes full responsibility and liability for offering the APIpie.ai API as part of your products - Clearly outlines your purpose and methods for the collection, storage, use, disclosure, and transfer of personal data in accordance with relevant privacy and data protection laws. - Remember, your commitment to privacy should be no less stringent than the requirements set out in our Privacy Policy and DPA. --- ### 7. Security Measures 7.1 You commit to implementing rigorous and robust security systems to protect your network, operating system, servers, databases, computers, user information, personal data, and other components constituting or supporting your Products integrating or using the APIpie.ai API ("your Systems"). If any of your Systems get compromised (e.g., through hacking, unauthorized use, or other security breaches), you should inform us immediately via email at . Upon notice, Neuronic AI may decide at its sole discretion whether to terminate your access to the APIpie.ai API. You acknowledge that such a security breach may negatively affect our reputation and the security of our other clients. Thus, you agree that termination due to a breach of your System is reasonable and necessary. Neuronic AI will not be held liable for any loss or damage related to such termination. --- ### 8. Ownership 8.1 All assets, rights, title, and interest in the Neuronic AI Brand, our API Documentation, APIpie.ai API, and our other Property, including but not limited to intellectual property rights, belong fully to Neuronic AI. We provide you with a limited license to use the API as set out in this API Agreement, and no other rights, title, interest, ownership, or property of any kind is transferred to you. Any inventions that you create incorporating our APIpie.ai API does not transfer any right of ownership or interest in our APIpie.ai API to you. You also agree to perform acts and execute documents (without compensation) as Neuronic AI may request from time to time to secure Neuronic AI's rights to our Property. --- ### 9. Termination 9.1 This API Agreement takes effect on the date that you agree to them or access or use the APIpie.ai API, whichever comes first, and will continue until terminated by Neuronic AI or you under the provisions of this API Agreement. 9.2 You may terminate this API Agreement by discontinuing your access to and/or use of our APIpie.ai API, and emailing us at of your intention to terminate the license granted hereunder and your use and access to the APIpie.ai API. 9.3 We may at any time vary, amend, change, suspend or discontinue provision of any of our Property, including the APIpie.ai API, suspend or terminate your use of the APIpie.ai API and/or the Neuronic AI Brand without notice or reasons to you. We may also restrict your access to or use of our APIpie.ai API if we determine that your access to or use of our APIpie.ai API may negatively impact our Products or Services. 9.4 Upon termination of this API Agreement: - 9.4.1 All rights and licenses granted to you hereunder will cease and terminate immediately; - 9.4.2 You should promptly destroy the API Documentation, Data and any other information procured from Neuronic AI or any of our Property that may be in your possession or control, and in any case within three (3) days of such termination; - 9.4.3 Unless we specifically grant our prior consent in writing, or unless otherwise stated in this API Agreement, you commit to promptly and permanently delete all Data and other information procured from or relating to Neuronic AI and any of our Property that you have stored in relation to your access to or use of the APIpie.ai API (including any **RAG Tuning** or memory data you may have stored). You agree to certify in writing the aforementioned destruction of stored Data and information should Neuronic AI request such certification from you. :br 9.5 All provisions under this API Agreement that must by their nature survive the termination of this API Agreement to fulfill their intention shall survive the termination of this API Agreement. --- ### 10. Your Warranties 10.1 You hereby represent and warrant to Neuronic AI that you (including any entity or other third party that you represent) have the ability to enter into and fully abide by this API Agreement and to use the APIpie.ai API in accordance with this API Agreement without violating any applicable laws or regulations, infringing any third party’s rights (including but not limited to intellectual property rights), and/or violating any other contracts you are bound by. 10.2 Without limiting the scope of the clauses under Section 11, you hereby warrant and undertake (and shall procure the same undertaking from any entity that you represent) not to initiate any legal actions or make any claims of any kind against Neuronic AI for any damages, expenses, or losses that you may suffer as a result of your use of (or inability to use) any of our Property (including but not limited to the APIpie.ai API). You acknowledge that we may, at any time and without notice or reference to you, amend this API Agreement or modify any aspect or function of the APIpie.ai API, which may adversely affect your usage of the APIpie.ai API and your associated business. This includes changes to the performance or cost of underlying LLM providers utilized by **AI Model Pooling** or **IMM**. In the event of any such adverse effects on your usage, you acknowledge and agree that your only recourse will be to terminate your use of the APIpie.ai API and this Agreement. You agree not to make any claims of any kind against Neuronic AI. For the avoidance of doubt, your continued access or use of the APIpie.ai API or any of our Property constitutes your acceptance of and agreement to all our aforesaid amendments and modifications. In the event of any dispute or adverse effect, you are encouraged to email us at . 10.3 Your use of and/or access to our Property (including the APIpie.ai API) is also subject to our Privacy Policy, User Agreement, and/or any other terms and conditions that we may implement from time to time in our discretion. In the event of any conflict between the various applicable agreements/terms and conditions, the terms of this API Agreement (based on its latest version) will take precedence in relation to your use of our APIpie.ai API. --- ### 11. Exclusions, Disclaimers and Limitation of Liability 11.1 Our Property, including but not limited to the APIpie.ai API, and all related components, data, documentation, and information are provided on an “AS IS” and “AS AVAILABLE” basis without any warranties of any kind, and Neuronic AI hereby expressly disclaims any and all warranties, whether express or implied, including, but not limited to, the implied warranties of merchantability, title, fitness for a particular purpose, continued accessibility, functionality, security, and non-infringement. You acknowledge that Neuronic AI does not warrant that our Property (including but not limited to the APIpie.ai API, data, and API documentation) will be uninterrupted, timely, updated, accurate, secure, error-free, omission-free, or virus-free, nor does it make any warranty as to the functions that may be performed or results that may be obtained from use of our Property (including but not limited to the APIpie.ai API). This includes the performance, reliability, and accuracy of outputs generated by third-party LLM providers utilized by the Service (even if routed via the **Integrity** or **RAG Tuning** features). 11.2 You (and any entity or other third party that you may represent) hereby agree that under no circumstances and under no legal theory (whether in contract, tort, equity, at law or otherwise) will Neuronic AI be liable to you or any third party for: - (I) any indirect, incidental, special, exemplary, consequential or punitive damages, including loss of profits, loss of sales, loss of opportunities, loss of business, loss of data, loss of reputation; or - (II) any amount for all claims cumulatively in excess of the fees actually paid by you to Neuronic AI in the six (6) months immediately preceding the earliest event giving rise to your claim or, if no fees have been paid by you to Neuronic AI, (the equivalent of) $100 USD; or - (III) any matter beyond our reasonable or foreseeable control. :br 11.3 You (and any entity or individual third party that you represent) agree to defend, hold harmless and indemnify Neuronic AI, its related companies, subsidiaries, affiliates, their respective officers, agents, employees, and suppliers, from and against any third-party claim arising from or in any way related to your or your users’/customers’ use of any of your Products, any of our Property (including but not limited to APIpie.ai API or Data), use of Neuronic AI Brand, or breach of any provision of this API Agreement, our Privacy Policy and/or User Agreement, including any liability or expense arising from all claims, losses, damages (actual and consequential), suits, judgments, settlement fees, and legal fees. --- ### 12. Governing Law and Jurisdiction This API Agreement and any claim, cause of action or dispute arising out of or related to this API Agreement shall be governed by the laws of Naples Florida, United States - regardless of any conflict of laws theory. Accordingly, you agree to submit to the exclusive jurisdiction of the Courts of Naples Florida. Notwithstanding the aforegoing, you agree that Neuronic AI shall be fully entitled to apply for equitable relief or injunctive remedies against you in any jurisdiction. --- ### 13. General 13.1 If any provision of this API Agreement is found to be illegal, void, or unenforceable, the said provision shall be modified so as to render it enforceable to the maximum extent possible in order to effect the intention of the provision; if such provision cannot be so modified, it shall be deleted and the remaining provisions of this API Agreement will continue in full force and effect. 13.2 The governing language of this API Agreement, API Documentation, the Site, all content, information and documents howsoever made available by Neuronic AI is English. In the event that Neuronic AI provides you with a translation of the English language version of this API Agreement, API Documentation, the Site or any other document, you understand that such translation is provided purely for your convenience and is not binding. 13.3 For any notices or service of legal documents, you agree that we may notify you via postings on the Site or via the email address that you may have provided to us. Neuronic AI accepts service of process by mail or courier at the email and address as stipulated in clause 13.9 below. Any notices that you provide without compliance with this clause 13.3 and clause 13.9 below will have no legal effect. 13.4 This API Agreement, our Privacy Policy, User Agreement and/or any documents incorporated into this API Agreement by reference, shall constitute the entire agreement between Neuronic AI and you regarding our Property and supersedes all prior agreements and understandings (whether via email correspondence or otherwise), whether written or oral, or whether established by custom, practice, policy or precedent, with respect to the subject matter of this API Agreement. 13.5 Any delays in or failure to act in relation to a breach of any provision of this API Agreement by you or others does not waive our right to act with respect to that breach or subsequent similar or other breaches. 13.6 Under no circumstances shall you seek or be entitled to rescission, injunctive or other equitable relief, or to enjoin or restrain the operation of the Site, any of our Services or our Property (including but not limited to the APIpie.ai API). 13.7 You are not permitted to assign, transfer, delegate or subcontract any rights or obligations hereunder this API Agreement without the prior written consent of Neuronic AI. However, you agree that Neuronic AI shall be entitled at any time to howsoever assign, transfer, delegate or subcontract any of its rights and/or obligations hereunder this API Agreement without any reference or notice to you. 13.8 No person or entity who is not a party to this API Agreement shall be entitled to rely on or enforce any provision herein this API Agreement; provided, however, that any affiliates of Neuronic AI (including but not limited to its licensor, if any) may rely on and enforce this Agreement. 13.9 You can contact us with your queries, comments and feedback by email to , however, for official notices and/or service of legal documents, your email must be followed by registered post or courier of such notices/legal documents to:- Neuronic AI - 1910 Pacific Ave Suite 2000 #1731, Dallas, Texas, 75201 # Data Processing Addendum (DPA) *Last updated on November 14, 2025* This Data Processing Addendum ("**DPA**") forms part of the Terms of Service or other written or electronic agreement (the "**Agreement**") between **Neuronic AI**, ("**Processor**", "**we**", "**us**", "**our**") and the entity or individual accepting this DPA ("**Customer**", "**Controller**"). This DPA governs Processor’s Processing of Personal Data on behalf of Customer. --- ## 1. Definitions **"Personal Data"** means any information relating to an identified or identifiable natural person as defined under GDPR, CCPA/CPRA, or other applicable data protection laws. **"Processing"** means any operation performed on Personal Data, including collection, storage, retrieval, transmission, or deletion. **"Sub-processor"** means any third party engaged by Processor to assist in Processing Personal Data. **"Services"** means the APIpie / Neuronic AI platform, model orchestration system, search services, memory services, vector services, and related tools used by Customer. **"Applicable Data Protection Laws"** include GDPR, CCPA/CPRA, and any similar laws. --- ## 2. Roles of the Parties Customer is the **Controller** of Personal Data. Processor acts as Customer’s **Processor** and will process Personal Data only to provide the Services and according to Customer’s documented instructions. Processor may engage approved **Sub-processors** as listed in Section 10. --- ## 3. Nature and Purpose of Processing Processor processes Personal Data for the following purposes: - Executing Customer API requests across supported LLM providers. - Routing, transforming, and enriching data as instructed by Customer. - Providing optional memory, RAG, vector, and stateful conversation features. - Performing search, scrape, and internet augmentation when requested by Customer. - Maintaining service logs required for operational integrity, billing, usage analytics, and security. - Supporting Customer-managed long-term collections for RAGTune or bot knowledge bases. Processor will not use Personal Data for training machine learning models unless explicitly documented by a Sub-processor and only when Customer intentionally selects such a model. --- ## 4. Categories of Data Subjects and Personal Data ### 4.1 Data Subjects May include Customer’s end-users, employees, clients, partners, or any individuals whose data Customer submits. ### 4.2 Types of Personal Data Personal Data processed may include: - Text or documents submitted by Customer - User identifiers for authentication - Prompt content submitted by Customer - Optional long-term memory content Processor does **not** retain prompt bodies for stateless API calls. --- ## 5. Processor Obligations Processor shall: - Process Personal Data **only** as instructed by Customer. - Maintain industry-standard security controls, encryption, and isolation. - Ensure data is encrypted in transit and at rest. - Restrict access to authorized personnel trained in data protection. - Delete Personal Data according to retention rules defined in Section 8. - Provide reasonable assistance with Data Subject Requests. - Notify Customer of any confirmed Personal Data breach without undue delay. - Maintain a record of Processing activities as required by law. --- ## 6. Customer Obligations Customer shall: - Ensure it has a lawful basis for Processing Personal Data via the Services. - Not submit unlawful or prohibited data categories without proper authorization. - Configure retention settings and model selection responsibly. - Avoid uploading unnecessary or excessive Personal Data. - Identify and disclose to Processor any special regulatory requirements. --- ## 7. Security Measures Processor implements: - Kubernetes-based isolated runtime using ephemeral compute. - In-memory Redis state storage (non-persistent) for short-lived conversation state. - Private Qdrant vector database under Processor’s full control for Customer memory and RAGTune. - Optional Pinecone vector database as a Sub-processor. - Routine deletion jobs to purge expired memory vectors. - Strict no-logging for prompt bodies. - Search anonymization (via Presidio) when Customer enables it. --- ## 8. Data Retention and Deletion ### 8.1 Stateless API Requests - No prompt bodies are retained. - Only usage metadata (model, token count, latency, cost) is retained. - Sub-processors may retain data for up to **30 days** solely for abuse detection. ### 8.2 Optional Memory Services When memory is enabled, Processor stores: - Conversation state (up to 24 hours unless Customer sets shorter expiry) - Long-term memory vectors (expiry configurable by Customer) Redis state is **in-memory only**. Vectors are stored only on Processor-controlled encrypted disks. ### 8.3 RAGTune Document Collections - Documents uploaded for persistent retrieval are stored until Customer deletes them. - Intended for support bots, knowledge bases, and enterprise content. ### 8.4 Deletion Requests - Customer may delete data at any time via API or request. - Processor runs scheduled deletion jobs ensuring expired data is removed. --- ## 9. International Data Transfers Processor and Sub-processors may transfer data internationally. Where required, Processor relies on: - Standard Contractual Clauses (SCCs) - Adequacy decisions - Other legal transfer mechanisms Customer may request the current list of regions and mechanisms. --- ## 10. Sub-processors Processor uses the following Sub-processors: ### 10.1 Primary LLM Providers - OpenAI - Anthropic - DeepSeek - OpenRouter - Together AI - DeepInfra - Eden AI ### 10.2 Vector & Memory Providers - Private Qdrant deployment (Processor-controlled) - Pinecone (optional, only if Customer enables) ### 10.3 Search & Scrape Providers - Google (search) - BrightData (search/scrape) - Valyu AI (primary internet agent) All Sub-processors operate under **no-training**, **restricted retention**, and **security compliance** as publicly documented. Processor will update Customer of material changes to Sub-processors. --- ## 11. Data Subject Rights Processor will assist Customer in responding to: - Access requests - Correction requests - Deletion requests - Objections - Data portability requests All such actions require Customer validation and instruction. --- ## 12. Audits and Compliance Customer may request information about Processor’s security measures. Formal audits may be conducted subject to reasonable controls to protect Processor’s platform and other customers. --- ## 13. Liability Each party’s liability under this DPA shall follow the limitations set forth in the Agreement. Processor is not liable for Customer’s use of non-compliant model providers or Customer’s failure to configure data retention or privacy settings appropriately. --- ## 14. Termination This DPA remains in effect as long as Processor processes Personal Data on behalf of Customer. Upon termination, Processor will delete or return all Personal Data except where retention is required by law. --- ## 15. Governing Law This DPA is governed by the same jurisdiction and governing law specified in the Agreement. --- ## 16. Acceptance By signing up for APIpie / Neuronic AI Services or continuing to use them, Customer agrees to the terms of this DPA. --- **END OF DPA** # Privacy Policy: Your Rights and Data Protection *Effective on April 1, 2022, Updated on November 14, 2025* This privacy policy ("Policy") describes how **Neuronic AI Inc d/b/a APIpie AI** ("Neuronic AI", "we", "us" or "our") collects, protects, and uses the personally identifiable information ("Personal Information") you ("User", "you" or "your") may provide on the APIpie.ai platform and any of its products or services (collectively, "**Services**"). This Policy specifically addresses how we manage your account data and, critically, how we process the data you submit to our AI platform via API calls (**User Content**). For specific details on our obligations when processing customer data on your behalf, please refer to our **Data Processing Addendum (DPA)**, located at [Data Processing Addendum (DPA)](https://apipie.ai/dpa). --- ### 1. Automatic Collection of Information (Usage Data) When you interact with the Services, our servers automatically record information related to that usage. This data may include: - **Request Metadata:** Your device's IP address, browser type/version, operating system, and access times. - **Service Metrics (Usage Logs):** Details on every API query (as provided by our **Usage** feature), including the specific model used, token counts, response times, costs, and unique user identifiers or source IPs. This statistical information is not otherwise aggregated in such a way that would identify any specific individual's User Content outside of your account. --- ### 2. Collection of Personal Information (Account Data) You can visit the marketing Website without providing Personal Information. However, to use the APIpie Services (e.g., to create an account and access the API), you will be asked to provide certain Personal Information. This typically includes your **email address** and **name**. We receive and store any information you knowingly provide to us when you create an account, publish content, make a purchase, or fill any online forms on the Services. When required, this information may include your email address, name, or other Personal Information. You can choose not to provide us with your Personal Information, but then you may not be able to take advantage of some of the Services' features. Users who are uncertain about what information is mandatory are welcome to contact us. --- ### 3. Collection and Handling of User Content (Prompt Data) **User Content** refers to the inputs (prompts, documents, files, images) and outputs generated when you use our core features like **Completions**, **RAG Tuning**, **Image Generation**, and **Voice Synthesis**. - **Stateless Processing:** For standard, stateless API calls, we generally implement a **strict no-logging policy for prompt bodies**. We process the data in memory to route it to the appropriate external model (**Models Route**, **Preferred Routing**, **AI Model Pooling**) and then discard the prompt body. Only Usage Logs (Section 1) are retained. - **Stateful Services (Memory/RAG):** If you enable stateful features (**RAG Tuning** or **Optional Memory Services**), the User Content you upload (documents, vector embeddings) or generate (conversation state) is stored securely on our infrastructure (or an authorized Sub-processor like Pinecone) and retained as long as required by your account settings or until you manually delete it. - **Anonymization Services:** We offer optional **parsing and pseudo-anonymization** services using technologies like Presidio and Tika. If you set `anon: true` in your prompt or use the direct anonymization service, we process the User Content to redact or anonymize potential Personal Data *before* transmitting it to any external entity or LLM provider. --- ### 4. Sharing and Disclosure of User Content (The Role of Sub-processors) Using the APIpie platform necessarily involves transmitting your User Content to third-party AI models and services to execute your requests. - **LLM Providers:** To deliver features like **Completions**, **Integrity**, and **Tools Support**, we must share your User Content with the specific external Large Language Model (LLM) or generative AI provider you or our **Intelligent Model Management (IMM)** system selects (e.g., OpenAI, Anthropic, ElevenLabs, DeepSeek). - **Vector/Search Providers:** To deliver **RAG Tuning** or augmented generation, User Content may be shared with or stored on vector and search sub-processors (e.g., Google, BrightData, or Pinecone, if enabled). We rely on the **data processing and no-training commitments** provided by our Sub-processors. **The full, current list of Sub-processors is detailed in the Data Processing Addendum (DPA).** By using our Services, you consent to the transfer of User Content to these Sub-processors for the purpose of fulfilling your API requests. --- ### 5. Managing Personal Information You are able to access, add to, update and delete certain Personal Information about you. The information you can view, update, and delete may change as the Services change. When you update information, however, we may maintain a copy of the unrevised information in our records. Some information may remain in our private records after your deletion of such information from your account. We will retain and use your Personal Information for the period necessary to comply with our legal obligations, resolve disputes, and enforce our agreements, unless a longer retention period is required or permitted by law. We may use any aggregated data derived from or incorporating your Personal Information after you update or delete it, but not in a manner that would identify you personally. Once the retention period expires, Personal Information shall be deleted. Therefore, the right to access, the right to erasure, the right to rectification, and the right to data portability cannot be enforced after the expiration of the retention period. --- ### 6. Use and Processing of Collected Information Any of the information we collect from you may be used to personalize your experience; improve our Services; improve customer service and respond to queries and emails of our customers; process transactions; send notification emails such as password reminders, updates, etc.; run and operate our Services. Information collected automatically is used only to identify potential cases of abuse and establish statistical information regarding Services usage. This statistical information is not otherwise aggregated in such a way that would identify any particular user of the system. We may process Personal Information related to you if one of the following applies: (i) You have given your consent for one or more specific purposes. Note that under some legislations we may be allowed to process information until you object to such processing (by opting out), without having to rely on consent or any other of the following legal bases. This, however, does not apply whenever the processing of Personal Information is subject to European data protection law; (ii) Provision of information is necessary for the performance of an agreement with you and/or for any pre-contractual obligations thereof; (iii) Processing is necessary for compliance with a legal obligation to which you are subject; (iv) Processing is related to a task that is carried out in the public interest or in the exercise of official authority vested in us; (v) Processing is necessary for the purposes of the legitimate interests pursued by us or by a third party. In any case, we will be happy to clarify the specific legal basis that applies to the processing, and in particular whether the provision of Personal Information is a statutory or contractual requirement, or a requirement necessary to enter into a contract. --- ### 7. Information Transfer and Storage Depending on your location, data transfers may involve transferring and storing your information in a country other than your own. You are entitled to learn about the legal basis of information transfers to a country outside the European Union or to any international organization governed by public international law or set up by two or more countries, such as the UN, and about the security measures taken by us to safeguard your information. If any such transfer takes place, you can find out more by checking the relevant sections of this document or inquire with us using the information provided in the contact section. We rely on legal transfer mechanisms, such as Standard Contractual Clauses (SCCs), where required for international transfers, as detailed in our DPA. --- ### 8. The Rights of Users You may exercise certain rights regarding your information processed by us. In particular, you have the right to do the following: (i) Withdraw consent where you have previously given your consent to the processing of your information; (ii) Object to the processing of your information if the processing is carried out on a legal basis other than consent; (iii) Learn if information is being processed by us, obtain disclosure regarding certain aspects of the processing, and obtain a copy of the information undergoing processing; (iv) Verify the accuracy of your information and ask for it to be updated or corrected; (v) Restrict the processing of your information, under certain circumstances; (vi) Obtain the erasure of your Personal Information from us, under certain circumstances; (vii) Receive your information in a structured, commonly used and machine-readable format and, if technically feasible, to have it transmitted to another controller without any hindrance. This provision is applicable provided that your information is processed by automated means and that the processing is based on your consent, on a contract which you are part of, or on pre-contractual obligations thereof. --- ### 9. The Right to Object to Processing Where Personal Information is processed for the public interest, in the exercise of an official authority vested in us, or for the purposes of the legitimate interests pursued by us, you may object to such processing by providing a ground related to your particular situation to justify the objection. However, should your Personal Information be processed for direct marketing purposes, you can object to that processing at any time without providing any justification. --- ### 10. How to Exercise These Rights Any requests to exercise User rights can be directed to Neuronic AI through the contact details provided in this document. These requests can be exercised free of charge and will be addressed by Neuronic AI as early as possible. --- ### 11. California Privacy Rights In addition to the rights explained in this Privacy Policy, California residents who provide Personal Information (as defined in the statute) to obtain products or services for personal, family, or household use are entitled to request and obtain from us, once a calendar year, information about the Personal Information we shared, if any, with other businesses for marketing uses. If applicable, this information would include the categories of Personal Information and the names and addresses of those businesses with which we shared such personal information for the immediately prior calendar year (e.g. requests made in the current year will receive information about the prior year). To obtain this information please contact us. --- ### 12. Billing and Payments In case of services requiring payment, we request credit card or other payment account information, which will be used solely for processing payments. Your purchase transaction data is stored only as long as is necessary to complete your purchase transaction. After that is complete, your purchase transaction information is deleted. All direct payment gateways adhere to the latest security standards as managed by the PCI Security Standards Council, which is a joint effort of brands like Visa, MasterCard, American Express and Discover. Sensitive and private data exchange happens over a SSL secured communication channel and is encrypted and protected with digital signatures. --- ### 13. Privacy of Children We do not knowingly collect any Personal Information from children under the age of 13. If you are under the age of 13, please do not submit any Personal Information through our Website or Service. We encourage parents and legal guardians to monitor their children's Internet usage and to help enforce this Policy by instructing their children never to provide Personal Information through our Website or Service without their permission. If you have reason to believe that a child under the age of 13 has provided Personal Information to us through our Website or Service, please contact us. You must also be old enough to consent to the processing of your personal data in your country (in some countries we may allow your parent or guardian to do so on your behalf). --- ### 14. Advertisement We may display online advertisements and we may share aggregated and non-identifying information about our customers that we collect through the registration process or through online surveys and promotions with certain advertisers. We do not share personally identifiable information about individual customers with advertisers. In some instances, we may use this aggregated and non-identifying information to deliver tailored advertisements to the intended audience. --- ### 15. Affiliates We may disclose information about you to our affiliates for the purpose of being able to offer you related or additional services. Any information relating to you that we provide to our affiliates will be treated by those affiliates in accordance with the terms of this Privacy Policy. --- ### 16. Links to Other Websites Our Services contain links to other websites that are not owned or controlled by us. Please be aware that we are not responsible for the privacy practices of such other websites or third-parties. We encourage you to be aware when you leave our Services and to read the privacy statements of each and every website that may collect Personal Information. --- ### 17. Changes and Amendments We reserve the right to modify this Privacy Policy relating to the Services at any time, effective upon posting of an updated version of this Privacy Policy on the Website. When we do, we will revise the updated date at the top of this page. Continued use of the Services after any such changes shall constitute your consent to such changes. --- ### 18. Acceptance of This Policy You acknowledge that you have read this Policy and agree to all its terms and conditions. By using the Services you agree to be bound by this Policy. If you do not agree to abide by the terms of this Policy, you are not authorized to use or access the Services. --- ### 19. Contacting Us If you have any questions about this Policy, please contact us. # Integration Overview ![APIpie](https://apipie.ai/img/docs/integrations/overview.svg) Welcome to the APIpie Integration Hub! This comprehensive overview showcases all available integrations organized by category. Whether you're building AI agents, creating chat applications, developing coding tools, or designing workflow automations, you'll find the perfect integration to power your project with APIpie's robust API. ## 🤖 Agent Frameworks Build sophisticated AI agents and multi-agent systems with these powerful frameworks: ### Core Frameworks - **[LangChain](https://apipie.ai/docs/Integrations/Agent-Frameworks/Langchain)** - Industry-standard framework for building LLM applications - **[LangGraph](https://apipie.ai/docs/Integrations/Agent-Frameworks/LangGraph)** - Graph-based agent orchestration from LangChain - **[CrewAI](https://apipie.ai/docs/Integrations/Agent-Frameworks/CrewAI)** - Multi-agent collaboration framework - **[AutoGen](https://apipie.ai/docs/Integrations/Agent-Frameworks/AutoGen)** - Multi-agent conversation framework by Microsoft - **[Llama-Index](https://apipie.ai/docs/Integrations/Agent-Frameworks/Llama-Index)** - Data framework for LLM applications ### Specialized Frameworks - **[DSPy](https://apipie.ai/docs/Integrations/Agent-Frameworks/DSPy)** - Programming framework for optimizing language model prompts - **[Composio](https://apipie.ai/docs/Integrations/Agent-Frameworks/Composio)** - Integration platform for AI agents - **[PydanticAI](https://apipie.ai/docs/Integrations/Agent-Frameworks/PydanticAI)** - Type-safe agent framework using Pydantic - **[Semantic-Kernel](https://apipie.ai/docs/Integrations/Agent-Frameworks/Semantic-Kernel)** - Microsoft's SDK for AI orchestration ### Development Platforms - **[HuggingFace](https://apipie.ai/docs/Integrations/Agent-Frameworks/HuggingFace)** - Open-source ML platform and model hub - **[Pixeltable](https://apipie.ai/docs/Integrations/Agent-Frameworks/Pixeltable)** - Data platform for AI applications - **[Vercel-AI](https://apipie.ai/docs/Integrations/Agent-Frameworks/Vercel-AI)** - AI SDK for building AI-powered applications - **[Smolagents](https://apipie.ai/docs/Integrations/Agent-Frameworks/Smolagents)** - Lightweight agent framework ### Enterprise Solutions - **[Google-ADK](https://apipie.ai/docs/Integrations/Agent-Frameworks/Google-ADK)** - Google Agent Development Kit - **[OpenAI-Agents](https://apipie.ai/docs/Integrations/Agent-Frameworks/OpenAI-Agents)** - OpenAI's agent development platform - **[AutoGPT](https://apipie.ai/docs/Integrations/Agent-Frameworks/AutoGPT)** - Autonomous AI agent framework - **[Agno](https://apipie.ai/docs/Integrations/Agent-Frameworks/Agno)** - Agent development framework --- ## 💬 Chat Agents Enhance conversational AI experiences with these popular chat platforms: ### Enterprise Chat Solutions - **[LibreChat](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat)** - Enhanced ChatGPT clone with multi-model support - **[OpenWebUI](https://apipie.ai/docs/Integrations/Chat-Agents/OpenwebUI)** - User-friendly WebUI for LLMs with offline capabilities - **[AnythingLLM](https://apipie.ai/docs/Integrations/Chat-Agents/AnythingLLM)** - Comprehensive AI application with RAG and agent capabilities ### Specialized Chat Interfaces - **[big-AGI](https://apipie.ai/docs/Integrations/Chat-Agents/big-AGI)** - Generative AI suite with advanced AGI functions - **[BetterChatGPT](https://apipie.ai/docs/Integrations/Chat-Agents/BetterChatGPT)** - Advanced UI for ChatGPT with enhanced features - **[ANSE](https://apipie.ai/docs/Integrations/Chat-Agents/ANSE)** - Multi-model platform optimized for various AI models - **[NextChat](https://apipie.ai/docs/Integrations/Chat-Agents/NextChat)** - Cross-platform ChatGPT/Gemini UI --- ## 💻 Coding Tools Supercharge your development workflow with AI-powered coding assistants: - **[Cline](https://apipie.ai/docs/Integrations/Coding/Cline)** - AI coding assistant for VS Code - **[Codex-CLI](https://apipie.ai/docs/Integrations/Coding/Codex-CLI)** - Command-line AI coding interface --- ## 🔧 Other Integrations Additional tools and platforms that enhance your AI development experience: - **[Helicone](https://apipie.ai/docs/Integrations/Other/Helicone)** - LLM observability and monitoring platform - **[LiteLLM](https://apipie.ai/docs/Integrations/Other/LiteLLM)** - Universal LLM API interface - **[NoApp](https://apipie.ai/docs/Integrations/Other/NoApp)** - AI App library --- ## 🔄 Workflows Automate and orchestrate your AI processes with workflow platforms: - **[Pipedream](https://apipie.ai/docs/Integrations/Workflows/Pipedream)** - no code / low code Serverless integration and workflow automation platform --- ## Getting Started Each integration comes with: - ✅ **Step-by-step setup guide** - Easy-to-follow instructions - ⚙️ **Configuration examples** - Ready-to-use code snippets - 🔗 **Live demos** - Try before you integrate - 📚 **Best practices** - Tips for optimal performance - 🛠️ **Troubleshooting** - Common issues and solutions ## Integration Benefits By integrating with APIpie, you gain access to: - **🌍 Multi-Model Support** - Access 100+ AI models from top providers - **💰 Cost Optimization** - Transparent pricing and usage controls - **⚡ High Performance** - Fast, reliable API endpoints - **🔒 Enterprise Security** - SOC 2 compliant infrastructure - **📊 Advanced Analytics** - Detailed usage insights and monitoring - **🔄 Easy Migration** - OpenAI-compatible endpoints for seamless switching ## Get Your Project Listed We welcome new integrations! If you've built a tool, platform, or framework that works with APIpie, we'd love to feature it here. **Requirements:** - Working integration with APIpie API - Documentation and setup guide - Active maintenance and support **To submit your integration:** 1. Join our [Discord Community](https://discord.gg/hs82THc9Tw){rel=""nofollow""} 2. Share your integration details 3. Provide documentation and examples 4. Get featured in our integration directory ## Need Help? - 📚 **Documentation** - Comprehensive guides for each integration - 💬 **Discord Support** - Join our [community](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for real-time help - 📧 **Direct Support** - Contact our integration team - 🎥 **Video Tutorials** - Visual guides for complex setups Start building today and unlock the full potential of AI with APIpie's extensive integration ecosystem! # AnythingLLM: Step-by-Step Configuration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![AnythingLLM](https://apipie.ai/img/docs/AnythingLLM.png){height="125" width="497"} :: This guide will walk you through the process of integrating AnythingLLM with APIpie to leverage the power of multiple AI models and enhance your chatbot experience. ## Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Install AnythingLLM - Follow the [AnythingLLM Installation Guide](https://docs.useanything.com/installation-desktop/overview){rel=""nofollow""} to install either locally or via Docker ### 5. Access the AnythingLLM Settings - Configure the LLM Preference either on initial setup or by changing the settings after configured. - Find your prefered LLM on the [APIpie.ai Dashboard](https://apipie.ai/dashboard/){rel=""nofollow""}. - Select APIpie as your prefered provider. - Paste your generated API key into the designated field. - Configure the appropite model name. #### Initial Setup: ::div{.docs-image-row} ![Initial Setup](https://apipie.ai/img/docs/integrations/AnythingLLM/Initial-Settings.png) :: #### Changing Settings: ::div{.docs-image-row} ![Settings](https://apipie.ai/img/docs/integrations/AnythingLLM/Settings.png) :: ### 6. Set the AnythingLLM Agent LLM #### Agent Setup: - From the main screen select the settings option on the configured workplace - Change to the "Agent Configuration" tab. - Select "APIpie" as the provider - Select the desired Agent AI model. ::div{.docs-image-row} ![Agent Setup](https://apipie.ai/img/docs/integrations/AnythingLLM/Workplace-Agent.png) :: ### 7. Start Chatting - With the integration complete, you can now enjoy AI-powered conversations in AnythingLLM. ::div{.docs-image-row} ![Desktop](https://apipie.ai/img/docs/integrations/AnythingLLM/AnythingLLM.png){height="502" width="845"} :: ![Desktop](https://apipie.ai/img/docs/integrations/AnythingLLM/AnythingLLM.png){width="100%"} ## Caveats - Currently AnythingLLM is yet support APIpie TTS, STT, or Embedding. - If this changes or if assistance is required to onboard more APIpie features please reach out on our [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} ## Additional Resources - **AnythingLLM Website**: Explore more features and capabilities of AnythingLLM. [Visit the website](https://useanything.com/){rel=""nofollow""} - **AnythingLLM GitHub**: Access the source code and contribute to the AnythingLLM project. [Visit GitHub](https://github.com/Mintplex-Labs/anything-llm){rel=""nofollow""} - **AnythingLLM Documentation**: Dive deeper into AnythingLLM's functionalities and configurations. [Access the docs](https://docs.useanything.com/){rel=""nofollow""} - **Mintplex Labs**: Explore Mintplex Labs the team that brings us AnythingLLM [Mintplex Labs](https://mintplexlabs.com/){rel=""nofollow""} ## Connect with AnythingLLM - [LinkedIn](https://www.linkedin.com/company/mintplex-labs){rel=""nofollow""} - [YouTube](https://www.youtube.com/@mintplexlabs){rel=""nofollow""} - [Twitter/X](https://x.com/AnythingLLM){rel=""nofollow""} - [Discord](https://discord.gg/Dh4zSZCdsC){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # BetterChatGPT Setup: Easy Integration Guide & Tips ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![BetterChatGPT](https://apipie.ai/img/docs/BetterChatGPT.png){height="125" width="125"} :: This guide will walk you through the process of integrating BetterChatGPT with APIpie to leverage the power of multiple AI models and enhance your chatbot experience. ## Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Open BetterChatGPT Settings - Launch the [BetterChatGPT Online Demo](https://bettergpt.chat/){rel=""nofollow""} the API setting will open automatically on first load or it can be accessed via the "API" option in the bottom left corner, then select the "View advanced API configuration" option. ::div{.docs-image-row} ![Initial API Settings](https://apipie.ai/img/docs/integrations/BetterChatGPT/initial-API.png) :: ### 5. Enter Your API Key - Check the "Use custom API endpoint" checkbox - Set the "API Endpoint" to {rel=""nofollow""} - Paste your generated API key into the designated field. - Save the settings. ::div{.docs-image-row} ![Advanced API Settings](https://apipie.ai/img/docs/integrations/BetterChatGPT/ADV-API.png) :: ### 6. Change the BetterChatGPT LLM - From the main screen select the preconfigured model at the top of the screen. ::div{.docs-image-row} ![Change LLM](https://apipie.ai/img/docs/integrations/BetterChatGPT/Change-Model.png) :: - Configure the desired model name. - For information regarding your preferred LLM find it on the [APIpie.ai Dashboard](https://apipie.ai/dashboard/){rel=""nofollow""} - Set the "Max Tokens" and other settings as desired. - Save the settings. ::div{.docs-image-row} ![Set LLM](https://apipie.ai/img/docs/integrations/BetterChatGPT/Set-Model.png) :: ### 7. Start Chatting - With the integration complete, you can now enjoy AI-powered conversations in BetterChatGPT. ::div{.docs-image-row} ![Desktop](https://apipie.ai/img/docs/integrations/BetterChatGPT/Chat.png) :: ## Caveats - Currently BetterChatGPT is yet to support APIpie TTS, STT, Embedding, or Image. - If this changes or if assistance is required to onboard more APIpie features please reach out on our [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} ## Additional Resources - **BetterChatGPT Website**: Explore more features and capabilities of BetterChatGPT. [Visit the website](https://bettergpt.chat/){rel=""nofollow""} - **BetterChatGPT GitHub**: Access the source code, Documentaion, and contribute to the BetterChatGPT project. [Visit GitHub](https://github.com/ztjhz/BetterChatGPT){rel=""nofollow""} ## Connect with BetterChatGPT - [Discord](https://discord.gg/g3Qnwy4V6A){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # big-AGI Integration: Quick Configuration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![big-AGI](https://apipie.ai/img/docs/big-AGI.png){height="125" width="223"} :: This guide will walk you through the process of integrating big-AGI with APIpie to leverage the power of multiple AI models and enhance your chatbot experience. ## Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Open big-AGI Settings - Launch the [big-AGI Online Demo](https://get.big-agi.com/){rel=""nofollow""} after signing in select the "Preferences" menu in the bottom left corner. - Under the "Chat" Tab select the "Models" button. - Select "OpenAI" fromt he Serivce options. - Paste your generated API key into the designated field. - Select "Advanced" to show the advanced settings. - Set the "API Endpoint" to {rel=""nofollow""} ::div{.docs-image-row} ![API Settings](https://apipie.ai/img/docs/integrations/big-AGI/Models-API.png) :: - Refresh the Models at the botton by selecting the blue "Models" button. - Find the Models you wish to use in the list, you can then modify the settings for each model and set your defaults. - You can set separate "Chat", "Fast", or "Function" models depending on your preference. ::div{.docs-image-row} ![Model Settings](https://apipie.ai/img/docs/integrations/big-AGI/Model-settings.png) :: ### 5. Start Chatting - With the integration complete, you can now enjoy AI-powered conversations in big-AGI. ::div{.docs-image-row} ![Desktop](https://apipie.ai/img/docs/integrations/big-AGI/big-AGI.png){height="461" width="1000"} :: ## Caveats - Currently big-AGI is yet to support APIpie TTS, STT, Image. - If this changes or if assistance is required to onboard more APIpie features please reach out on our [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} ## Additional Resources - **big-AGI Website**: Explore more features and capabilities of big-AGI. [Visit the website](https://get.big-agi.com/){rel=""nofollow""} - **big-AGI News**: Stay up to date with big-AGI news. [Visit the website](https://get.big-agi.com/news){rel=""nofollow""} - **big-AGI GitHub**: Access the source code and contribute to the big-AGI project. [Visit GitHub](https://github.com/enricoros/big-agi){rel=""nofollow""} - **big-AGI Documentation**: Dive deeper into big-AGI's functionalities and configurations. [Access the docs](https://github.com/enricoros/big-AGI/blob/main/docs/README.md){rel=""nofollow""} - **big-AGI Demo**: Try out big-AGI before integrating it. [Online Demo](https://get.big-agi.com/){rel=""nofollow""} ## Connect with big-AGI - [Discord](https://discord.gg/MkH4qj2Jp9){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # LibreChat Configuration Guide: Quick Setup ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![LibreChat](https://apipie.ai/img/docs/librechat.png){height="125" width="125"} :: This guide will walk you through the process of integrating LibreChat with APIpie to leverage the power of multiple AI models and enhance your chatbot experience. ## Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Open LibreChat Settings - Launch the [LibreChat Online Demo](https://demo.librechat.cfd/){rel=""nofollow""} and dropdown the provider list and click the settings icon next to APIpie. ::div{.docs-image-row} ![Settings](https://apipie.ai/img/docs/integrations/librechat/apipie-settings.png) :: ### 5. Enter Your API Key - Paste your generated API key into the designated field and save the settings. ::div{.docs-image-row} ![Add Key](https://apipie.ai/img/docs/integrations/librechat/add-api-key.png) :: ### 6. Start Chatting - With the integration complete, you can now enjoy AI-powered conversations in LibreChat. #### Desktop ::div{.docs-image-row} ![Desktop](https://apipie.ai/img/docs/integrations/librechat/librechat-apipie.png){height="500" width="1106"} :: #### Mobile ::div{.docs-image-row} ![Mobile](https://apipie.ai/img/docs/integrations/librechat/librechat-mobile.png){height="500" width="231"} :: ## Caveats - Currently NextChat is yet to support APIpie TTS, STT, Image. - If this changes or if assistance is required to onboard more APIpie features please reach out on our [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} ## Additional Resources - **LibreChat Website**: Explore more features and capabilities of LibreChat. [Visit the website](https://librechat.ai/){rel=""nofollow""} - **LibreChat GitHub**: Access the source code and contribute to the LibreChat project. [Visit GitHub](https://github.librechat.ai/){rel=""nofollow""} - **LibreChat Documentation**: Dive deeper into LibreChat's functionalities and configurations. [Access the docs](https://www.librechat.ai/docs/configuration/librechat_yaml/ai_endpoints/apipie){rel=""nofollow""} - **LibreChat Demo**: Try out LibreChat before integrating it. [Online Demo](https://demo.librechat.cfd/){rel=""nofollow""} | [HuggingFace](https://hf.librechat.ai/){rel=""nofollow""} ## Connect with LibreChat - [LinkedIn](https://linkedin.librechat.ai/){rel=""nofollow""} - [YouTube](https://www.youtube.com/@LibreChat){rel=""nofollow""} - [Twitter](https://x.com/LibreChatAI){rel=""nofollow""} - [Discord](https://discord.librechat.ai/){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # NextChat Integration: Simple Setup Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![NextChat](https://apipie.ai/img/docs/NextChat.png){height="125" width="656"} :: This guide will walk you through the process of integrating NextChat with APIpie to leverage the power of multiple AI models and enhance your chatbot experience. ## Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Open NextChat Settings - Launch the [NextChat Online Demo](https://app.nextchat.dev/){rel=""nofollow""} then click the settings icon in the bottom left of the screen, then scroll until you see the API settings. - Check the "Custom Endpoint" checkbox. - Keep "Model Provider" as OpenAI. - Set the "OpenAI Endpoint" to {rel=""nofollow""} - Paste your generated API key into the designated field. - Choose your desired model. - Adjust models settings if desired. - Save your settings. ::div{.docs-image-row} ![API Settings](https://apipie.ai/img/docs/integrations/NextChat/API.png) :: ### 5. Start Chatting - With the integration complete, you can now enjoy AI-powered conversations in NextChat. ::div{.docs-image-row} ![Chat](https://apipie.ai/img/docs/integrations/NextChat/chat.png) :: ## Caveats - Currently NextChat is yet to support APIpie TTS, STT, Image. - If this changes or if assistance is required to onboard more APIpie features please reach out on our [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} ## Additional Resources - **NextChat Website**: Explore more features and capabilities of NextChat. [Visit the website](https://app.nextchat.dev/){rel=""nofollow""} - **NextChat GitHub**: Access the source code and contribute to the NextChat project. [Visit GitHub](https://github.com/ChatGPTNextWeb/ChatGPT-Next-Web){rel=""nofollow""} - **NextChat FAQ**: Dive deeper into NextChat's functionalities and configurations. [Access the FAQ](https://github.com/ChatGPTNextWeb/ChatGPT-Next-Web/blob/main/docs/faq-en.md){rel=""nofollow""} - **NextChat Demo**: Try out NextChat before integrating it. [Online Demo](https://app.nextchat.dev/){rel=""nofollow""} ## Connect with NextChat - [Twitter/X](https://twitter.com/NextChatDev){rel=""nofollow""} - [Discord](https://discord.gg/YCkeafCafC){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Complete OpenWebUI Setup: Configuration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![OpenWebUI](https://apipie.ai/img/docs/OpenWebUI.png){height="125" width="125"} :: This guide will walk you through the process of integrating OpenWebUI with APIpie to leverage the power of multiple AI models and enhance your chatbot experience. ## Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Install OpenWebUI - Follow the [OpenWebUI Installation Guide](https://docs.openwebui.com/getting-started/){rel=""nofollow""} to install either manually or via Docker ### 5. Access the OpenWebUI Settings #### Configure APIpie API: - Open the Settings within OpenWebUI and navigate to the Connections Tab - Enable OpenAI API - Set the "API Base URL" to {rel=""nofollow""} - Paste your generated API key into the designated field. - Verify the connection is successful - Save the settings. ::div{.docs-image-row} ![Configure API](https://apipie.ai/img/docs/integrations/OpenWebUI/API-config.png) :: #### Set Default LLM: - Navigate to the Users Tab - Set the Default Model with your preferred LLM - Save the settings. ::div{.docs-image-row} ![Set LLM](https://apipie.ai/img/docs/integrations/OpenWebUI/Set-LLM.png) :: ### 6. Start Chatting - With the integration complete, you can now enjoy AI-powered conversations in OpenWebUI. ::div{.docs-image-row} ![Desktop](https://apipie.ai/img/docs/integrations/OpenWebUI/OpenWebUI.png){height="453" width="867"} :: ## Caveats - Currently OpenWebUI is yet to support APIpie TTS, STT, Image, & Embedding. - If this changes or if assistance is required to onboard more APIpie features please reach out on our [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} ## Additional Resources - **OpenWebUI Website**: Explore more features and capabilities of OpenWebUI. [Visit the website](https://openwebui.com/){rel=""nofollow""} - **OpenWebUI GitHub**: Access the source code and contribute to the OpenWebUI project. [Visit GitHub](https://github.com/open-webui/open-webui){rel=""nofollow""} - **OpenWebUI Documentation**: Dive deeper into OpenWebUI's functionalities and configurations. [Access the docs](https://docs.openwebui.com/){rel=""nofollow""} ## Connect with OpenWebUI - [LinkedIn](https://www.linkedin.com/company/open-webui/){rel=""nofollow""} - [Twitter](https://twitter.com/OpenWebUI){rel=""nofollow""} - [Discord](https://discord.com/invite/5rJgQTnV4s){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Integrate ANSE with APIpie: Step-by-Step Setup Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){width="125"} ![ANSE](https://apipie.ai/img/docs/ANSE.png){width="286"} :: This guide will walk you through the process of integrating ANSE with APIpie to leverage the power of multiple AI models and enhance your chatbot experience. ## Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Open ANSE Settings - Launch the [ANSE Online Demo](https://anse.app/){rel=""nofollow""} then click the pencil icon to the right of the "OpenAI" settings at the right of the screen. - Choose your desired model. - Paste your generated API key into the designated field. - Set the "Base URL" to {rel=""nofollow""} ::div{.docs-image-row} ![API Settings](https://apipie.ai/img/docs/integrations/ANSE/API.png) :: ### 5. Start Chatting - With the integration complete, you can now enjoy AI-powered conversations in ANSE. ::div{.docs-image-row} ![Chat](https://apipie.ai/img/docs/integrations/ANSE/chat.png){style="margin-right: 20px;"} :: ## Caveats - Currently ANSE is yet to support APIpie TTS, STT, Image. - If this changes or if assistance is required to onboard more APIpie features please reach out on our [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} ## Additional Resources - **ANSE Website**: Explore more features and capabilities of ANSE. [Visit the website](https://anse.app/){rel=""nofollow""} - **ANSE GitHub**: Access the source code and contribute to the ANSE project. [Visit GitHub](https://github.com/anse-app/anse){rel=""nofollow""} - **ANSE Documentation**: Dive deeper into ANSE's functionalities and configurations. [Access the docs](https://docs.anse.app/){rel=""nofollow""} - **ANSE Demo**: Try out ANSE before integrating it. [Online Demo](https://anse.app/){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Codex CLI Integration Guide: APIpie Setup ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![OpenAI - Codex CLI ](https://apipie.ai/img/docs/openai-white.png){height="108" width="400"} :: This guide will walk you through integrating OpenAI Codex CLI with APIpie, enabling advanced AI-powered coding, automation, and reasoning directly from your terminal. ## What is Codex CLI? Codex CLI is a lightweight, open-source coding agent that runs in your terminal. It brings ChatGPT-level reasoning and automation to your development workflow, allowing you to: - Run code, manipulate files, and iterate under version control - Use multimodal input (text, screenshots, diagrams) - Operate in interactive or fully automated modes - Leverage advanced models for code, reasoning, and project automation By connecting Codex CLI to APIpie, you unlock access to a wide range of powerful models, enhanced context windows, and cost-efficient AI capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with Codex CLI. ### 4. Install Codex CLI Install globally via npm: ```shell npm install -g @openai/codex ``` Or with yarn: ```shell yarn global add @openai/codex ``` ### 5. Configure Codex CLI for APIpie Set your APIpie endpoint and API key as environment variables: ```shell export APIPIE_BASE_URL="https://apipie.ai/v1" export APIPIE_API_KEY="your-APIpie-key-here" ``` ::tip Add these lines to your shell profile (e.g., `~/.zshrc`, `~/.bashrc`) for persistence. :: You can now run Codex CLI with APIpie: ( you only need to use the flag "--provider APIPIE" to change codex across to APIPIE as a provider ) ```shell codex --provider APIPIE ``` Or with a prompt: ```shell codex --provider APIPIE "explain this codebase to me" ``` ### 6. (Optional) Advanced Configuration Codex CLI supports a config file at `~/.codex/config.yaml` where you can configure APIpie as your provider: ```yaml model: gpt-4o-mini # or any supported APIpie model provider: APIPIE providers: APIPIE: name: 'APIpie' baseURL: 'https://apipie.ai/v1' envKey: 'APIPIE_API_KEY' fullAutoErrorMode: ask-user ``` Alternatively, you can use JSON format at `~/.codex/config.json`: ```json { "model": "gpt-4o-mini", "provider": "APIPIE", "providers": { "APIPIE": { "name": "APIpie", "baseURL": "https://apipie.ai/v1", "envKey": "APIPIE_API_KEY" } }, "fullAutoErrorMode": "ask-user" } ``` You can also add custom instructions in `~/.codex/instructions.md` or use the new-style project docs in `~/.codex/AGENTS.md`. ::tip Browse available models on the [APIpie Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} to see which models you can access. :: ::div{.docs-image-row} ![OpenAI - Codex CLI ](https://apipie.ai/img/docs/integrations/Codex/demo.gif){width="800"} :: --- ## Key Features - **Zero Setup:** Just set your APIpie API key and endpoint—no extra configuration required. - **Full Auto-Approval:** Use `--approval-mode full-auto` for hands-free automation (sandboxed and safe). - **Multimodal Input:** Pass screenshots or diagrams for multimodal reasoning. - **Project Memory:** Codex CLI merges project docs and instructions for context-aware automation. - **Sandboxed Execution:** All commands run network-disabled and directory-sandboxed for security. - **Model Flexibility:** Access APIpie's latest models (e.g., o3, o4-mini, GPT-4.1, Grok, Claude, Llama, etc.) for coding, reasoning, and more. --- ## Example Workflows | What you type | What happens | | ---------------------------------------------------------------------- | ------------------------------------------------------------- | | `codex "Refactor the Dashboard component to React Hooks"` | Codex rewrites the component, runs tests, and shows the diff. | | `codex "Generate SQL migrations for adding a users table"` | Creates migration files and runs them in a sandboxed DB. | | `codex "Write unit tests for utils/date.ts"` | Generates and runs tests, iterating until they pass. | | `codex "Bulk-rename *.jpeg → *.jpg with git mv"` | Safely renames files and updates usages. | | `codex "Explain what this regex does: ^(?=.*[A-Z]).{8,}$"` | Outputs a step-by-step human explanation. | | `codex "Carefully review this repo, and propose 3 high impact PRs"` | Suggests impactful PRs in the current codebase. | | `codex "Look for vulnerabilities and create a security review report"` | Finds and explains security bugs. | --- ## Security & Permissions Codex CLI lets you control agent autonomy via the `--approval-mode` flag: - **Suggest (default):** Only reads files; all writes and shell commands require approval. - **Auto Edit:** Reads and writes files; shell commands require approval. - **Full Auto:** Reads/writes files and executes shell commands (all sandboxed). All actions are sandboxed and network-disabled for safety. --- ## Troubleshooting & FAQ - **Does Codex CLI work on Windows?**:br Yes, via [WSL2](https://learn.microsoft.com/en-us/windows/wsl/install){rel=""nofollow""}. Native support is available for macOS and Linux. - **Which models are supported?**:br Any model available via APIpie's OpenAI-compatible endpoint (e.g., o3, o4-mini, GPT-4.1, Grok, Claude, Llama, etc.) - **How do I persist my API key and endpoint?**:br Add the `export` lines to your shell profile. - **How do I view available APIpie models?**:br Visit the [APIpie Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} to see all models you can use with Codex CLI. For more, see the [Codex CLI GitHub](https://github.com/openai/codex){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Cursor Configuration Guide: Quick Setup ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![Cursor](https://apipie.ai/img/docs/integrations/Cursor/Cursor.png){height="125"} :: This guide will walk you through the process of integrating Cursor with APIpie to leverage the power of multiple AI models and enhance your AI coding experience. ::tip ##### Free Tier Cursor Users If you are on the Free Tier with Cursor and are running into the "Model XXX Does not work with your current plan or API key" issue then follow our [Cursor Free Tier Workaround Guide](https://apipie.ai/docs/blog/Cursors-Does-Not-Work-with-Your-Current-Plan-or-API-Key-Fix) :: ## What is Cursor? Cursor is an AI-powered code editor that revolutionizes the way developers write, edit, and understand code. Built on the foundation of Visual Studio Code, Cursor integrates advanced AI capabilities directly into your coding workflow, offering intelligent code completion, natural language code generation, and contextual assistance that understands your entire codebase. ## Steps ### 1. Create an APIpie Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Download and Install Cursor - Visit the official [Cursor website](https://cursor.com/home){rel=""nofollow""} - Download the appropriate version for your operating system (Windows, macOS, or Linux) - Install Cursor following the standard installation process for your platform ::div{.docs-image-row} ![Download Cursor](https://apipie.ai/img/docs/integrations/Cursor/cursor-download.png) :: ### 5. Configure APIpie in Cursor Settings - Open Cursor and navigate to Settings (Cursor > Settings > Cursor Settings) - Go to the "Models" section in the settings sidebar - Configure the following APIpie integration: #### Basic Configuration - Scroll down to the "OpenAI API Key" section - Toggle the switch to enable the OpenAI API Key functionality - Click on "Override OpenAI Base URL" and enter: `https://apipie.ai/v1` - Input your APIpie API key from step 3 ::div{.docs-image-row} ![APIpie Settings](https://apipie.ai/img/docs/integrations/Cursor/apipie-settings.png) :: #### Model Configuration - Disable all models - Click the Verify button to check your configuration - Add APIpie models by clicking "+ Add model" and entering model identifiers (e.g., `claude-3-5-sonnet-20241022`, `gpt-4o`) ::tip ##### Troubleshooting - You must disable any model not provided by your custom endpoint - APIpie :: ::div{.docs-image-row} ![APIpie Models](https://apipie.ai/img/docs/integrations/Cursor/apipie-models.png) :: ### 6. Start Coding with Cursor and APIpie - With the integration complete, you can now enjoy enhanced AI-powered coding assistance using Cursor with APIpie models - Start typing code and experience intelligent completions powered by APIpie's diverse model selection ::tip ##### Tips for Best Results - Ensure you have sufficient credits in your APIpie account for uninterrupted coding assistance - Disable other API providers in Cursor to avoid conflicts with APIpie integration :: ## Testing Your Integration 1. **Test Tab Autocomplete**: Start typing a function and see intelligent completions appear 2. **Try Agent Mode**: Press Ctrl+I and ask to create a new function or refactor existing code 3. **Use Inline Edit**: Select code, press Ctrl+K, and describe changes you want to make 4. **Chat Interface**: Open the chat panel and ask questions about your codebase ## Additional Resources - **Cursor Website**: Learn more about Cursor's features and capabilities. [Visit the website](https://cursor.com/features){rel=""nofollow""} - **Cursor Documentation**: Comprehensive guides and tutorials for using Cursor. [Access the docs](https://docs.cursor.com/){rel=""nofollow""} - **Cursor Community**: Connect with other Cursor users and get support. [Join the community](https://forum.cursor.com/){rel=""nofollow""} - **APIpie Observability Dashboard**: Monitor your API usage and manage your account. [Visit Observability](https://apipie.ai/profile/activity){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Cline Configuration Guide: Quick Setup ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![Cline](https://apipie.ai/img/docs/integrations/Cline/Cline.png){height="125" width="125"} :: This guide will walk you through the process of integrating Cline with APIpie to leverage the power of multiple AI models and enhance your AI assistant experience in VSCode. ## What is Cline? Cline is an AI assistant that can use your CLI and Editor. Thanks to advanced agentic coding capabilities, Cline can handle complex software development tasks step-by-step. With tools that let it create & edit files, explore large projects, use the browser, and execute terminal commands (after you grant permission), it can assist you in ways that go beyond code completion or tech support. ## Steps ### 1. Create an APIpie Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Install Cline Extension in VSCode - Open VSCode and navigate to the Extensions marketplace - Search for "Cline APIpie" and install the extension - Alternatively, you can install it directly from the [VS Marketplace](https://marketplace.visualstudio.com/items?itemName=NeuronicAI.cline-apipie){rel=""nofollow""} ::div{.docs-image-row} ![Install Extension](https://apipie.ai/img/docs/integrations/Cline/install-extension.png) :: ### 5. Configure APIpie in Cline with Advanced Features - Open Cline in VSCode - Click the Cline-APIpie icon in the Activity Bar - Click on the settings icon in the interface - Select "API Providers" and choose "APIpie" from the dropdown - Enter your APIpie API key in the designated field - Configure the following advanced features: #### Enable Memory Memory provides prompt caching functionality similar to Anthropic's caching, but works across all models. This allows you to: - Maintain conversation context when switching between different models - Reduce token usage by avoiding reprocessing of previous messages - Create consistent experiences across different API providers **Memory Configuration Options:** - **Session**: A unique identifier for the current conversation thread - **Expire**: Time in days before cached messages expire - **mem\_msgs**: Maximum number of messages to keep in memory For more information about memory read our [Integrated Model Memory (IMM)](https://apipie.ai/docs/features/imm) docs. **Important**: Always clear memory between separate tasks and different sessions using the "Clear Memory" button to prevent context contamination and ensure fresh responses. #### Enable Integrity Integrity is a powerful feature that helps prevent hallucinations in AI responses. When enabled, this feature: - Verifies information against known facts - Reduces the likelihood of generating inaccurate or fabricated information - Improves the reliability of responses for factual queries For more information about memory read our [Integrity](https://apipie.ai/docs/features/integrity) docs. ::div{.docs-image-row} ![APIpie Settings](https://apipie.ai/img/docs/integrations/Cline/apipie-settings.png) :: ### 6. Start Using Cline with APIpie - With the integration complete, you can now enjoy AI-powered assistance in VSCode using Cline with APIpie models ::div{.docs-image-row} ![Cline Interface](https://apipie.ai/img/docs/integrations/Cline/demo.gif) :: ## Key Features ### Run Commands in Terminal Cline can execute commands directly in your terminal and receive the output. This allows it to perform a wide range of tasks, from installing packages and running build scripts to deploying applications, managing databases, and executing tests. ::div{.docs-image-row} ![Terminal Commands](https://apipie.ai/img/docs/integrations/Cline/terminal-commands.png) :: ### Create and Edit Files Cline can create and edit files directly in your editor, presenting you a diff view of the changes. You can edit or revert Cline's changes directly in the diff view editor, or provide feedback in chat until you're satisfied with the result. ::div{.docs-image-row} ![File Editing](https://apipie.ai/img/docs/integrations/Cline/file-editing.png) :: ### Use the Browser Cline can launch a browser, click elements, type text, and scroll, capturing screenshots and console logs at each step. This allows for interactive debugging, end-to-end testing, and even general web use. ::div{.docs-image-row} ![Browser Use](https://apipie.ai/img/docs/integrations/Cline/browser-use.png) :: ### Custom Tools with MCP Thanks to the Model Context Protocol, Cline can extend its capabilities through custom tools. Just ask Cline to "add a tool" and it will handle everything, from creating a new MCP server to installing it into the extension. ::div{.docs-image-row} ![Custom Tools](https://apipie.ai/img/docs/integrations/Cline/custom-tools.png) :: ## Caveats - Make sure to review and approve all terminal commands before execution - For optimal performance, use Claude 3.5 Sonnet or higher models - If you encounter any issues with the integration, please reach out on our [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} ## Additional Resources - **Cline Website**: Explore more features and capabilities of Cline. [Visit the website](https://cline.bot/){rel=""nofollow""} - **Cline GitHub**: Access the source code and contribute to the Cline project. [Visit GitHub](https://github.com/cline/cline){rel=""nofollow""} - **Cline Documentation**: Dive deeper into Cline's functionalities and configurations. [Access the docs](https://docs.cline.bot/){rel=""nofollow""} - **Cline Discord**: Join the community for support and discussions. [Join Discord](https://discord.gg/cline){rel=""nofollow""} ## Connect with Cline - [Twitter](https://twitter.com/clinebot){rel=""nofollow""} - [Reddit](https://www.reddit.com/r/cline/){rel=""nofollow""} - [Feature Requests](https://feedback.cline.bot/){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Pipedream Integration: Setup Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![Pipedream](https://apipie.ai/img/docs/integrations/Pipedream/pd.jpg){height="125" width="125"} :: This guide will walk you through the process of integrating [Pipedream](https://pipedream.com/){rel=""nofollow""} with APIpie to leverage the power of multiple AI models and enhance your chatbot experience. ## Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Open Pipedream your Project - From within your desired Pipedream project add a new action where you desire and then search for APIpie. ::div{.docs-image-row} ![Search APIpie](https://apipie.ai/img/docs/integrations/Pipedream/Search_APIpie.png) :: - Select the desired action, either standard API Request, Node.js, or Python, this guide will follow the standard API Request. ::div{.docs-image-row} ![Action Choice](https://apipie.ai/img/docs/integrations/Pipedream/Action_Choice.png) :: ### 5. Connect Your APIpie API Key - Select "Connect an APIpie.ai Account" and then enter your API key into the designated field and save the settings. ::div{.docs-image-row} ![Connect APIpie](https://apipie.ai/img/docs/integrations/Pipedream/Connect_APIpie.png) :: - Enter your API key into the designated field and save the settings, or test conenction to ensure it's working correctly. ::div{.docs-image-row} ![Enter API Key](https://apipie.ai/img/docs/integrations/Pipedream/APIkey.png) :: ### 6. Configure API Request Details - You can then enter in your desired API Request details, for any clarification of details rever to the [APIpie Documentation](https://apipie.ai/docs/api/introduction){rel=""nofollow""}. ::div{.docs-image-row} ![Enter API Request Details](https://apipie.ai/img/docs/integrations/Pipedream/API_Request.png) :: You can now continue your Pipedream project with further actions as needed. ## Additional Resources - **Pipedream Website**: Explore more features and capabilities of Pipedream. [Visit the website](https://pipedream.com/){rel=""nofollow""} - **Pipedream GitHub**: Access the source code and contribute to the Pipedream project. [Visit GitHub](https://github.com/PipedreamHQ/pipedream){rel=""nofollow""} - **Pipedream Documentation**: Dive deeper into Pipedreams's functionalities and configurations. [Visit the Docs](https://pipedream.com/apps/apipie-ai){rel=""nofollow""} ## Connect with OpenWebUI - [LinkedIn](https://www.linkedin.com/company/pipedreamhq/){rel=""nofollow""} - [Twitter](https://twitter.com/pipedream){rel=""nofollow""} - [Support](https://pipedream.com/support){rel=""nofollow""} - [Community](https://pipedream.com/community/){rel=""nofollow""} ### Additional Support If you encounter any issues during the integration process, please reach out to us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # LiteLLM Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![LiteLLM](https://apipie.ai/img/docs/integrations/LiteLLM/LiteLLM.png){width="600"} :: This guide will walk you through integrating LiteLLM with APIpie, enabling you to access a wide range of language models through a unified OpenAI-compatible interface. ## What is [LiteLLM](https://litellm.ai/){rel=""nofollow""}? LiteLLM is a powerful library and proxy service that provides a unified interface for working with multiple LLM providers including: - **Unified API Interface**: Call any LLM API using the familiar OpenAI format - **Consistent Response Format**: Get standardized responses regardless of the underlying provider - **Fallback Routing**: Configure backup models in case of API failures or rate limits - **Cost Tracking & Budgeting**: Monitor spend across providers and set limits - **Proxy Server**: Deploy a gateway to manage all your LLM API calls By integrating LiteLLM with APIpie, you can leverage APIpie's powerful models while maintaining compatibility with other providers through a single consistent interface. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with LiteLLM. ### 4. Install LiteLLM ```bash pip install litellm ``` For the proxy server: ```bash pip install 'litellm[proxy]' ``` ### 5. Configure LiteLLM for APIpie There are two main ways to use LiteLLM with APIpie: #### A. Direct Integration (Python Library) ```python import os import litellm from litellm import completion # Set your APIpie API key os.environ["APIPIE_API_KEY"] = "your-apipie-api-key" # Use APIpie models with LiteLLM response = completion( model="apipie/gpt-4o-mini", # Use APIpie's gpt-4o-mini model messages=[{"role": "user", "content": "Hello, how are you?"}] ) print(response) ``` #### B. Configure LiteLLM Proxy with APIpie Create a `config.yaml` file: ```yaml model_list: - model_name: gpt-4 litellm_params: model: apipie/gpt-4 api_key: your-apipie-api-key api_base: 'https://apipie.ai/v1' - model_name: apipie/claude-3-sonnet litellm_params: model: claude-3-sonnet-20240229 api_key: your-apipie-api-key api_base: 'https://apipie.ai/v1' - model_name: apipie/* litellm_params: model: apipie/* api_key: your-apipie-api-key api_base: 'https://apipie.ai/v1' ``` Then start the proxy: ```bash litellm --config config.yaml ``` --- ## Key Features - **Provider-Agnostic Interface**: Switch between models from different providers without changing your code - **Fallback Routing**: Configure fallback models if your primary model fails - **Consistent Response Format**: Get standardized responses in OpenAI format - **Cost Management**: Track and limit spending across different providers - **Streaming Support**: Stream responses from all supported models - **Logging & Monitoring**: Track usage, errors, and performance --- ## Example Workflows | Application Type | What LiteLLM Helps You Build | | --------------------------- | -------------------------------------------------------- | | Multi-Provider Applications | Apps that can use any LLM provider through a single API | | High-Reliability Services | Systems with automatic fallback to backup models | | Cost-Optimized Applications | Apps that route to the most cost-efficient provider | | Enterprise LLM Gateways | Central API gateways with access controls and monitoring | | Multi-Model Agents | Agents that use specialized models for different tasks | --- ## Using LiteLLM with APIpie ### Basic Chat Completion ```python import os from litellm import completion # Set your APIpie API key os.environ["APIPIE_API_KEY"] = "your-apipie-api-key" # Use any APIpie model response = completion( model="apipie/gpt-4o", # APIpie's GPT-4o messages=[ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Write a short poem about AI."} ] ) print(response.choices[0].message.content) ``` ### Streaming Responses ```python import os from litellm import completion # Set your APIpie API key os.environ["APIPIE_API_KEY"] = "your-apipie-api-key" # Stream the response response = completion( model="apipie/gpt-4o", messages=[{"role": "user", "content": "Write a short story about space exploration."}], stream=True ) # Process the streaming response for chunk in response: content = chunk.choices[0].delta.content if content: print(content, end="", flush=True) ``` ### Model Fallbacks with Router ```python import os from litellm import Router # Set your APIpie API key os.environ["APIPIE_API_KEY"] = "your-apipie-api-key" # Configure models with fallbacks model_list = [ { "model_name": "gpt-4", "litellm_params": { "model": "apipie/gpt-4", "api_key": os.environ["APIPIE_API_KEY"], "api_base": "https://apipie.ai/v1" } }, { "model_name": "apipie/claude-alternative", "litellm_params": { "model": "apipie/claude-3-sonnet-20240229", "api_key": os.environ["APIPIE_API_KEY"], "api_base": "https://apipie.ai/v1" } }, { "model_name": "apipie/*", "litellm_params": { "model": "apipie/*", "api_key": os.environ["APIPIE_API_KEY"], "api_base": "https://apipie.ai/v1" } } ] # Initialize the router with fallback options router = Router( model_list=model_list, fallbacks=[ {"gpt-4": ["claude-alternative"]} # If gpt-4 fails, try claude ] ) # Use the router for completions response = router.completion( model="gpt-4", messages=[{"role": "user", "content": "Explain quantum computing in simple terms."}] ) print(response.choices[0].message.content) ``` ### Async Completions ```python import os import asyncio from litellm import acompletion # Set your APIpie API key os.environ["APIPIE_API_KEY"] = "your-apipie-api-key" async def get_completion(): response = await acompletion( model="apipie/gpt-4o", messages=[{"role": "user", "content": "What are the benefits of renewable energy?"}] ) return response.choices[0].message.content # Run the async function result = asyncio.run(get_completion()) print(result) ``` ### Using Embeddings ```python import os from litellm import embedding # Set your APIpie API key os.environ["APIPIE_API_KEY"] = "your-apipie-api-key" # Generate embeddings response = embedding( model="apipie/text-embedding-3-large", input=["Renewable energy is the future of sustainable living."] ) print(f"Embedding dimension: {len(response.data[0].embedding)}") print(f"First few values: {response.data[0].embedding[:5]}") ``` --- ## Setting up the LiteLLM Proxy with APIpie The LiteLLM proxy serves as a gateway between your applications and various LLM providers, including APIpie. ### Step 1: Create a Configuration File Create a `config.yaml` file: ```yaml model_list: - model_name: gpt-4 litellm_params: model: apipie/gpt-4 api_key: your-apipie-api-key api_base: 'https://apipie.ai/v1' - model_name: apipie/claude-3-sonnet litellm_params: model: claude-3-sonnet-20240229 api_key: your-apipie-api-key api_base: 'https://apipie.ai/v1' - model_name: apipie/* litellm_params: model: apipie/* api_key: your-apipie-api-key api_base: 'https://apipie.ai/v1' # Optional configurations router_settings: routing_strategy: 'simple-shuffle' # or "usage-based", "latency-based" # Set up the proxy server litellm_settings: success_callback: ['prometheus'] # For metrics tracking drop_params: true # Drop unsupported parameters # Set up API keys for proxy users (optional) api_keys: - key: 'sk-1234' aliases: ['team-1'] metadata: team: 'research' spend: 0 max_budget: 100 # $100 budget ``` ### Step 2: Start the Proxy Server ```bash litellm --config config.yaml ``` The proxy will start on `http://localhost:4000` by default. ### Step 3: Use the Proxy in Your Applications ```python import openai client = openai.OpenAI( api_key="sk-1234", # Your proxy API key base_url="http://localhost:4000/v1" # Your proxy URL ) response = client.chat.completions.create( model="gpt-4", # This will be routed to APIpie/gpt-4 based on your config messages=[{"role": "user", "content": "Explain the theory of relativity."}] ) print(response.choices[0].message.content) ``` --- ## Troubleshooting & FAQ - **How do I handle environment variables securely?**:br Store your API keys in environment variables and avoid hardcoding them in your code. For production, use a secure secret management system. - **How can I monitor costs across providers?**:br LiteLLM's proxy server includes spend tracking functionality. You can also integrate with observability tools like Langfuse or Helicone. - **Can I set rate limits on API usage?**:br Yes, the LiteLLM proxy allows setting rate limits per user, team, or model. - **Does LiteLLM support function calling or tool use?**:br Yes, when using OpenAI-compatible models through APIpie that support function calling, those capabilities are preserved. - **How do I handle provider-specific parameters?**:br Use the `litellm_params` field in your configuration to specify provider-specific parameters. - **Can I deploy the proxy on cloud platforms?**:br Yes, the LiteLLM proxy can be deployed on platforms like Render, Railway, AWS, GCP, or Azure. For more information, see the [LiteLLM documentation](https://docs.litellm.ai/docs/){rel=""nofollow""} or the [GitHub repository](https://github.com/BerriAI/litellm){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} or [LiteLLM Discord](https://discord.gg/wuPM9dRgDw){rel=""nofollow""} for assistance. # Helicone Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![Helicone](https://apipie.ai/img/docs/integrations/Helicone/dark.svg){height="125"} :: This guide will walk you through integrating Helicone with APIpie, enabling comprehensive observability, monitoring, and analytics for your LLM applications with just a few lines of code. ## What is [Helicone](https://helicone.ai/){rel=""nofollow""}? Helicone is an all-in-one, open-source LLM developer platform that provides: - **Observability**: Track and monitor all your LLM requests with detailed logs and metrics - **Analytics**: Gain insights into usage patterns, costs, latency, and other key metrics - **Tracing**: Visualize the execution path of your LLM applications and agent systems - **Prompt Management**: Version and experiment with prompts using production data - **Evaluations**: Run automated evaluations on your LLM outputs - **Security & Compliance**: Enterprise-ready with SOC 2 and GDPR compliance By connecting Helicone with APIpie, you can monitor all your LLM interactions across different models with minimal overhead, gaining valuable insights to optimize performance and costs. --- ## Integration Steps ### 1. Create a Helicone Account - **Register here:** [Helicone Registration](https://us.helicone.ai/signup){rel=""nofollow""} - Helicone offers a generous free tier with 10,000 requests per month. ### 2. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 3. Add Credit to APIpie - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 4. Generate API Keys - **APIpie API Key:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - **Helicone API Key:** Navigate to your Helicone dashboard and generate an API key ### 5. Install Required Libraries ```bash # For Python pip install openai # or pip install together # If you're using the Together AI client ``` For JavaScript/TypeScript: ```bash # For Node.js npm install openai # or yarn add openai # or pnpm add openai ``` --- ## Key Features - **One-line Integration**: Add observability to your LLM applications with minimal code changes - **Real-time Monitoring**: Track request volume, costs, latency, and more in real-time - **Session & Trace Analysis**: Analyze complex agent systems and multi-step LLM workflows - **Cost Management**: Track spending across models and optimize for cost-efficiency - **Prompt Management**: Version control and optimize your prompts - **Caching**: Reduce costs and improve performance with built-in caching --- ## Example Workflows | Application Type | What Helicone Helps You Monitor | | --------------------------- | --------------------------------------------------------- | | Chatbot Applications | Track conversation flows, user satisfaction, and costs | | Agent Systems | Visualize complex agent interactions and tool usage | | RAG Implementations | Monitor retrieval quality and overall system performance | | Production LLM Applications | Ensure reliability, manage costs, and track KPIs | | Prompt Engineering | Compare different prompt versions and their effectiveness | --- ## Using Helicone with APIpie ### Basic Python Integration (Base URL Method) ```python import os import openai # Configure the client with Helicone proxy and APIpie key client = openai.OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), # Your APIpie API key base_url=f"https://oai.hconeai.com/v1/{os.environ.get('HELICONE_API_KEY')}" ) # Make requests as normal - Helicone will log them automatically response = client.chat.completions.create( model="gpt-4o-mini", # Use any model available on APIpie messages=[ {"role": "user", "content": "What are some fun things to do in London?"} ] ) print(response.choices[0].message.content) ``` ### Python Integration (Headers Method - More Secure) ```python import os import openai # Configure the client with Helicone headers client = openai.OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), # Your APIpie API key base_url="https://apipie.ai/v1", # APIpie endpoint default_headers={ "Helicone-Auth": f"Bearer {os.environ.get('HELICONE_API_KEY')}" } ) # Make requests as normal - Helicone will log them automatically response = client.chat.completions.create( model="gpt-4o-mini", # Use any model available on APIpie messages=[ {"role": "user", "content": "What are some fun things to do in London?"} ] ) print(response.choices[0].message.content) ``` ### JavaScript/TypeScript Integration ```typescript import OpenAI from 'openai'; // Configure the client with Helicone proxy const openai = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, // Your APIpie API key baseURL: `https://oai.hconeai.com/v1/${process.env.HELICONE_API_KEY}`, }); // Or use headers for more secure environments const openaiWithHeaders = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, // Your APIpie API key baseURL: 'https://apipie.ai/v1', // APIpie endpoint defaultHeaders: { 'Helicone-Auth': `Bearer ${process.env.HELICONE_API_KEY}`, }, }); // Make requests as normal - Helicone will log them automatically async function getCompletion() { const response = await openai.chat.completions.create({ model: 'gpt-4o-mini', // Use any model available on APIpie messages: [{ role: 'user', content: 'What are some fun things to do in London?' }], }); console.log(response.choices[0].message.content); } getCompletion(); ``` ### Adding Custom Properties to Requests Helicone allows you to add custom properties to your requests for better filtering and analysis: ```python import os import openai client = openai.OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), # Your APIpie API key base_url="https://apipie.ai/v1", # APIpie endpoint default_headers={ "Helicone-Auth": f"Bearer {os.environ.get('HELICONE_API_KEY')}", "Helicone-Property-User-Id": "user_123", # Add custom user ID "Helicone-Property-Session-Id": "session_abc", # Add session tracking "Helicone-Property-App-Version": "1.2.3", # Track app version } ) response = client.chat.completions.create( model="gpt-4o-mini", messages=[ {"role": "user", "content": "What are some fun things to do in London?"} ] ) print(response.choices[0].message.content) ``` ### Caching Responses for Cost Optimization Enable caching to avoid redundant API calls and reduce costs: ```python import os import openai client = openai.OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1", default_headers={ "Helicone-Auth": f"Bearer {os.environ.get('HELICONE_API_KEY')}", "Helicone-Cache-Enabled": "true" # Enable caching } ) # Make the same request multiple times - only the first will hit the API for i in range(3): response = client.chat.completions.create( model="gpt-4o-mini", messages=[ {"role": "user", "content": "What is the capital of France?"} ] ) print(f"Request {i+1}: {response.choices[0].message.content}") ``` ### Tracking Feedback Collect user feedback to evaluate model performance: ```python import os import requests def log_feedback(request_id, rating, comment=None): """Log user feedback for a specific request.""" url = "https://api.hconeai.com/v1/feedback" headers = { "Authorization": f"Bearer {os.environ.get('HELICONE_API_KEY')}", "Content-Type": "application/json" } payload = { "request_id": request_id, "rating": rating, # 1 for bad 5 for good "comment": comment } response = requests.post(url, headers=headers, json=payload) return response.json() # Use after getting a response request_id = response.id # Get this from the Helicone response log_feedback(request_id, 5, "Perfect answer!") ``` --- ## Viewing Your Data in Helicone After integrating Helicone with APIpie, you can access your dashboards and analyze your data: 1. Log in to your [Helicone Dashboard](https://us.helicone.ai/dashboard){rel=""nofollow""} 2. View key metrics like request volume, costs, and latency 3. Explore traces and sessions to understand user interactions 4. Analyze prompt effectiveness and model performance 5. Set up custom dashboards and alerts --- ## Troubleshooting & FAQ - **I'm not seeing any data in Helicone**:br Ensure your API keys are correct and that you're making requests through the Helicone proxy or with the proper headers. - **How does Helicone affect my API latency?**:br Helicone adds minimal latency typically less than 10 to your requests when using their cloud version. - **Is my data secure with Helicone?**:br Yes, Helicone is SOC 2 and GDPR compliant. For maximum security, consider self-hosting Helicone using their Docker or Helm charts. - **How can I filter requests in Helicone?**:br Use custom properties in your headers to add metadata to requests, which can then be filtered in the dashboard. - **Can I export my data from Helicone?**:br Yes, Helicone provides APIs for data export, allowing you to integrate with your existing data pipelines or analytics tools. - **Is there a self-hosted option?**:br Yes, Helicone is open-source and can be self-hosted using Docker or Helm. See their [documentation](https://docs.helicone.ai/getting-started/self-deploy-docker){rel=""nofollow""} for details. For more information, see the [Helicone documentation](https://docs.helicone.ai/){rel=""nofollow""} or their [GitHub repository](https://github.com/Helicone/helicone){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} or [Helicone Discord](https://discord.gg/zsSTcH2qhG){rel=""nofollow""} for assistance. # AG2 (AutoGen) Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![AG2 (AutoGen)](https://apipie.ai/img/docs/integrations/AG2/ag2-white.svg){height="125"} :: This guide will walk you through integrating AG2 (formerly AutoGen) with APIpie, enabling you to build powerful multi-agent AI applications with access to a wide range of language models. ## What is [AG2](https://ag2.ai/){rel=""nofollow""}? AG2 (formerly AutoGen) is an open-source framework for building and orchestrating AI agents. It provides a flexible and powerful system for: - **Multi-agent conversations** where agents can communicate and collaborate - **Human-in-the-loop workflows** for oversight and feedback - **Tool use** to extend agents with custom functionality - **Code execution** for solving complex problems - **Customizable agent behaviors** through system messages and specialized configurations By connecting AG2 with APIpie, you gain access to a wide range of powerful models and features while leveraging AG2's robust agent orchestration capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with AG2. ### 4. Install AG2 Install AG2 (AutoGen) using pip: ```bash pip install ag2[openai] # Or use the alias if you prefer pip install autogen[openai] ``` ### 5. Configure AG2 for APIpie Create a configuration file that points to APIpie's API: ```python import os from autogen import LLMConfig # Option 1: Using environment variables os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Option 2: Using configuration file (recommended) config_list = [ { "model": "gpt-4o-mini", # You can use any model available on APIpie "api_key": "your-apipie-api-key", "base_url": "https://apipie.ai/v1", } ] # Create an LLMConfig object llm_config = LLMConfig.from_config_dict(config_list[0]) ``` You can also save your configuration to a JSON file for better organization: ```python import json import os from autogen import LLMConfig # Create a configuration file config = [ { "model": "gpt-4o-mini", "api_key": "your-apipie-api-key", "base_url": "https://apipie.ai/v1", } ] # Save to a file (make sure to add it to .gitignore) with open("oai_config.json", "w") as f: json.dump(config, f) # Load from file llm_config = LLMConfig.from_json(path="oai_config.json") ``` --- ## Key Features - **Multi-Agent Systems**: Create collaborative agent networks that communicate to solve complex tasks - **Customizable Agents**: Define specialized agents with unique roles, knowledge, and capabilities - **Tool Integration**: Extend agents with custom functions and external APIs - **Code Generation & Execution**: Generate and run code to solve technical problems - **Human-in-the-Loop**: Include human feedback and oversight at key decision points --- ## Example Workflows | Application Type | What AG2 Helps You Build | | ------------------------ | -------------------------------------------------------------- | | Conversational Agents | AI assistants that can engage in natural dialogue | | Problem-Solving Teams | Groups of specialized agents that collaborate on complex tasks | | Coding Assistants | Agents that can write, debug, and execute code | | Research Aids | Agents that can gather, analyze, and synthesize information | | Decision Support Systems | AI workflows that help humans make informed decisions | --- ## Using AG2 with APIpie ### Basic Multi-Agent Setup ```python import os from autogen import AssistantAgent, UserProxyAgent, LLMConfig # Configure APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Or load from configuration llm_config = LLMConfig(api_type="openai", model="gpt-4o-mini") # Create an assistant that uses APIpie with llm_config: assistant = AssistantAgent( name="assistant", system_message="You are a helpful AI assistant specializing in data analysis." ) # Create a user proxy agent that can execute code user_proxy = UserProxyAgent( name="user_proxy", code_execution_config={"work_dir": "coding", "use_docker": False} ) # Start a conversation user_proxy.initiate_chat( assistant, message="Analyze the following data and create a visualization: [1, 5, 3, 8, 2, 7, 4]" ) ``` ### Using Multiple Specialized Agents ```python import os from autogen import AssistantAgent, UserProxyAgent, LLMConfig, GroupChat, GroupChatManager # Configure APIpie llm_config = LLMConfig(api_type="openai", model="gpt-4o-mini") with llm_config: # Create specialized agents planner = AssistantAgent( name="planner", system_message="You break down complex tasks into manageable steps. Be concise and clear.", ) researcher = AssistantAgent( name="researcher", system_message="You find and provide information needed to complete tasks. Focus on reliable sources.", ) coder = AssistantAgent( name="coder", system_message="You write code to solve problems. Explain your code clearly.", ) critic = AssistantAgent( name="critic", system_message="You review solutions and suggest improvements. Be constructive.", ) # User proxy for human input and code execution user_proxy = UserProxyAgent( name="user_proxy", code_execution_config={"work_dir": "coding", "use_docker": False}, is_termination_msg=lambda msg: "TASK COMPLETE" in msg.get("content", ""), ) # Create a group chat groupchat = GroupChat( agents=[user_proxy, planner, researcher, coder, critic], messages=[], max_round=15, ) # Create a manager to orchestrate the conversation manager = GroupChatManager( groupchat=groupchat, llm_config=llm_config, ) # Start the conversation user_proxy.initiate_chat( manager, message="Create a Python script that pulls weather data for New York City and displays it as a chart." ) ``` ### Implementing Tool Usage ```python import os from typing import List, Dict, Any from autogen import register_function, AssistantAgent, UserProxyAgent, LLMConfig # Configure APIpie llm_config = LLMConfig(api_type="openai", model="gpt-4o-mini") # Define a custom tool def search_products(query: str, max_results: int = 5) -> List[Dict[str, Any]]: """ Search for products matching the query. Args: query: The search query string. max_results: Maximum number of results to return. Returns: A list of product dictionaries with name, price, and rating. """ # In a real application, this would call your API # This is just a mock example mock_data = [ {"name": "Smartphone X", "price": 899.99, "rating": 4.5}, {"name": "Wireless Earbuds", "price": 149.99, "rating": 4.3}, {"name": "Laptop Pro", "price": 1299.99, "rating": 4.7}, {"name": "Smart Watch", "price": 249.99, "rating": 4.2}, {"name": "Bluetooth Speaker", "price": 79.99, "rating": 4.4}, ] filtered = [p for p in mock_data if query.lower() in p["name"].lower()] return filtered[:max_results] # Create an assistant with tool use capability with llm_config: assistant = AssistantAgent( name="shopping_assistant", system_message="You help users find products. Use the search_products tool when needed.", ) # Create a user proxy user_proxy = UserProxyAgent( name="user", human_input_mode="TERMINATE", code_execution_config={"work_dir": "shopping", "use_docker": False}, ) # Register the custom tool register_function( search_products, caller=assistant, executor=user_proxy, description="Search for products matching a query string", ) # Start the conversation user_proxy.initiate_chat( assistant, message="I'm looking for wireless audio devices. Can you help me find some options?" ) ``` --- ## Troubleshooting & FAQ - **Which models are supported?**:br Any model available via APIpie's OpenAI-compatible endpoint. - **How do I handle environment variables securely?**:br Store your API key in environment variables or in a configuration file that's added to your .gitignore. Never commit API keys to repositories. - **Can I mix different LLMs within the same agent system?**:br Yes, you can create different LLMConfig objects for different agents, allowing them to use different models based on their specific needs. - **What if I need to handle requests with large context windows?**:br Choose models on APIpie that support larger context windows, like gpt-4-turbo or claude-3-opus. - **How can I optimize costs when using AG2 with APIpie?** Use smaller/cheaper models for simple tasks and reserve more powerful models for complex reasoning. APIpie's routing capabilities can help optimize this automatically. For more information, see the [AG2 documentation](https://docs.ag2.ai/){rel=""nofollow""} or the [GitHub repository](https://github.com/ag2ai/ag2){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # AutoGPT Classic: Step-by-Step Configuration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![AutoGPT](https://apipie.ai/img/docs/autogpt.png){height="125" width="276"} :: This guide will walk you through the process of setting up your environment to work with APIpie and integrate AutoGPT functionalities effectively. ## Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4 Configure Your .env File - Find the `.env` file in your project's directory: `autogpts/autogpt/`. If the file doesn't exist, rename `.env.template` to `.env`. - Open the `.env` file in your preferred text editor. ### 5. Set Your API Key - Add or edit the following line with your new API key: `OPENAI_API_KEY=your_new_api_key_here` ### 6. Update the Base URL - On line 68, remove the `#` to uncomment the line and update the base URL to `OPENAI_API_BASE_URL=https://apipie.ai/v1` ### 7. Save and Close the File - Save the changes and close your text editor. ## Additional Resources - **AutoGPT User Guide**: Learn more about how to use AutoGPT effectively with their detailed guide. [View the guide here](https://docs.agpt.co/classic/usage/){rel=""nofollow""}. - **In-depth AutoGPT Documentation**: Dive deeper into the functionalities and configurations of AutoGPT with their comprehensive documentation. [Access the docs here](https://autogptdocs.com/){rel=""nofollow""}. - **AutoGPT News**: Keep up to date with the latest updates and features from AutoGPT. [Visit the news site](https://news.agpt.co/){rel=""nofollow""}. - **AutoGPT GitHub Repository**: Explore the source code and contribute to the AutoGPT project on GitHub. [Visit GitHub](https://github.com/Significant-Gravitas/AutoGPT/tree/master){rel=""nofollow""}. ## Connect with AutoGPT - [Twitter](https://twitter.com/Auto_GPT){rel=""nofollow""} - [Discord](https://discord.gg/autogpt){rel=""nofollow""} ### Additional Support If this documentation is outdated, or if you encounter any issues, please get in touch with us on [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""}. # Composio Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![Composio](https://apipie.ai/img/docs/integrations/Composio/composio_white_font.svg){height="125"} :: This guide will walk you through integrating Composio with APIpie, enabling your AI applications to leverage over 250+ tools and services through a unified interface. ## What is [Composio](https://composio.dev/){rel=""nofollow""}? Composio is a production-ready toolset for AI agents that provides a unified interface to over 250+ external tools and services. Key features include: - **Extensive Tool Library**: Access to tools like GitHub, Notion, Linear, Gmail, Slack, Hubspot, and more - **OS Operations**: File management, shell commands, code analysis, and other system-level operations - **Search Capabilities**: Integrations with Google, Perplexity, Tavily, Exa, and other search engines - **Managed Authentication**: Support for multiple auth protocols (OAuth, API Keys, Basic JWT) - **Optimized Tool Calling**: Up to 40% improved accuracy through specialized tool design - **Framework Compatibility**: Works with LLM frameworks like LangChain, CrewAI, AutoGen, and more By connecting Composio with APIpie, you can enable your AI agents to interact with external services and tools while leveraging APIpie's powerful model selection and routing capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with Composio. ### 4. Install Composio Install Composio and the OpenAI plugin: ```bash pip install composio-core composio-openai ``` For JavaScript/TypeScript projects: ```bash npm install composio-core # or yarn add composio-core # or pnpm add composio-core ``` ### 5. Configure Composio with APIpie Set up your environment variables: ```bash # Get your API keys from respective platforms export COMPOSIO_API_KEY=your-composio-api-key export OPENAI_API_KEY=your-apipie-api-key export OPENAI_BASE_URL=https://apipie.ai/v1 ``` --- ## Key Features - **Pre-built API Connectors**: 250+ integrations ready to use without custom development - **Authentication Management**: Simplified auth handling for all connected tools - **Tool Calling Optimization**: Enhanced accuracy for AI tool usage - **Multi-Framework Support**: Works with various AI frameworks and LLM providers - **Unified Interface**: Consistent API for all external tools - **Extensibility**: Ability to create custom tool definitions --- ## Example Workflows | Application Type | What Composio Helps You Build | | --------------------- | --------------------------------------------------------------- | | Software Integrations | AI agents that interact with GitHub, Notion, Slack, etc. | | Email & Communication | Tools that read, write, and manage emails and messages | | Data Analysis | Applications that fetch, analyze, and visualize external data | | Research Assistants | AI that can search and compile information from various sources | | Development Workflows | Code review, repository management, and CI/CD automation | --- ## Using Composio with APIpie ### Basic Python Integration This example shows how to use Composio with APIpie to star a GitHub repository: ```python from openai import OpenAI from composio_openai import ComposioToolSet, App, Action import os # Set up environment variables (can also be done in .env file) os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_BASE_URL"] = "https://apipie.ai/v1" os.environ["COMPOSIO_API_KEY"] = "your-composio-api-key" # Initialize OpenAI client with APIpie configuration openai_client = OpenAI() # Initialize the Composio Tool Set composio_tool_set = ComposioToolSet() # Connect your GitHub account (only needed once) # Run 'composio add github' in terminal and follow the auth flow # Get GitHub tools that are pre-configured actions = composio_tool_set.get_actions( actions=[Action.GITHUB_STAR_A_REPOSITORY_FOR_THE_AUTHENTICATED_USER] ) # Define your task my_task = "Star a repo composiodev/composio on GitHub" # Create an assistant with the GitHub tools assistant = openai_client.beta.assistants.create( name="GitHub Assistant", instructions="You are a helpful assistant that can interact with GitHub", model="gpt-4o-mini", # Use any model supported by APIpie tools=actions, ) # Create a thread thread = openai_client.beta.threads.create() # Add user message to thread message = openai_client.beta.threads.messages.create( thread_id=thread.id, role="user", content=my_task ) # Execute the assistant with tool capabilities run = openai_client.beta.threads.runs.create( thread_id=thread.id, assistant_id=assistant.id ) # Handle the tool calls and process the responses response = composio_tool_set.wait_and_handle_assistant_tool_calls( client=openai_client, run=run, thread=thread, ) print(response) ``` ### Basic JavaScript Integration ```javascript import { OpenAIToolSet } from 'composio-core'; import OpenAI from 'openai'; // Configure APIpie as the OpenAI provider const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY, baseURL: 'https://apipie.ai/v1', }); // Initialize Composio toolset const toolset = new OpenAIToolSet({ apiKey: process.env.COMPOSIO_API_KEY, }); // Connect your GitHub account (only needed once) // Run 'composio add github' in terminal and follow the auth flow // Get GitHub tools const tools = await toolset.getTools({ actions: ['GITHUB_STAR_A_REPOSITORY_FOR_THE_AUTHENTICATED_USER'], }); // Create GitHub assistant const githubAssistant = await openai.beta.assistants.create({ name: 'Github Assistant', instructions: "You're a GitHub Assistant, you can perform operations on GitHub", tools: tools, model: 'gpt-4o-mini', // Use any model supported by APIpie }); // Create thread const thread = await openai.beta.threads.create(); // Create run const run = await openai.beta.threads.runs.create(thread.id, { assistant_id: githubAssistant.id, instructions: "Star the repository 'composiohq/composio'", tools: tools, model: 'gpt-4o-mini', stream: false, }); // Handle tool calls const response = await toolset.waitAndHandleAssistantToolCalls(openai, run, thread); console.log(response); ``` ### Multi-Tool Workflow Example This example demonstrates using multiple tools together: ```python from openai import OpenAI from composio_openai import ComposioToolSet, Action import os # Configure APIpie and Composio os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_BASE_URL"] = "https://apipie.ai/v1" os.environ["COMPOSIO_API_KEY"] = "your-composio-api-key" # Initialize clients openai_client = OpenAI() composio_tool_set = ComposioToolSet() # Get multiple tool categories actions = composio_tool_set.get_actions( actions=[ # GitHub actions Action.GITHUB_LIST_REPOSITORIES_FOR_THE_AUTHENTICATED_USER, Action.GITHUB_GET_A_REPOSITORY, # Notion actions Action.NOTION_RETRIEVE_A_PAGE, Action.NOTION_CREATE_A_PAGE, # Google search action Action.GOOGLE_SEARCH ] ) # Create research assistant that can use all these tools assistant = openai_client.beta.assistants.create( name="Research Assistant", instructions=""" You are a research assistant that can: 1. Search for information on Google 2. Check GitHub repositories for code examples 3. Create and retrieve Notion pages to store findings Help users find information and organize it effectively. """, model="gpt-4o", # Use any model supported by APIpie tools=actions, ) # Create a thread thread = openai_client.beta.threads.create() # Add a complex research task message = openai_client.beta.threads.messages.create( thread_id=thread.id, role="user", content="Research the latest developments in AI agents, find some GitHub repositories with examples, and create a Notion page summarizing your findings." ) # Execute the assistant run = openai_client.beta.threads.runs.create( thread_id=thread.id, assistant_id=assistant.id ) # Process all tool calls response = composio_tool_set.wait_and_handle_assistant_tool_calls( client=openai_client, run=run, thread=thread, ) print(response) ``` --- ## Troubleshooting & FAQ - **How do I authenticate with external services?**:br Run `composio add ` in your terminal to start the authentication flow for any service (e.g., `composio add github`). - **Which models are supported?**:br Any model available via APIpie's OpenAI-compatible endpoint can be used with Composio. - **How do I handle environment variables securely?**:br Store your API keys in environment variables or use environment management tools like `python-dotenv`. Never commit API keys to repositories. - **Can I use Composio with other frameworks?**:br Yes, Composio supports multiple frameworks including LangChain, CrewAI, AutoGen, and others. - **How do I add custom tools?**:br Composio supports custom tool definitions. See the [Composio documentation](https://docs.composio.dev/framework){rel=""nofollow""} for details. For more information, see the [Composio documentation](https://docs.composio.dev/){rel=""nofollow""} or the [GitHub repository](https://github.com/composiohq/composio){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # CrewAI Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"}s ![CrewAI](https://apipie.ai/img/docs/integrations/CrewAI/crewai_logo.png){height="125"} :: This guide will walk you through integrating CrewAI with APIpie, enabling you to build powerful multi-agent systems and orchestrate complex AI workflows with access to a wide range of language models. ## What is [CrewAI](https://crewai.com/){rel=""nofollow""}? CrewAI is a fast and flexible open-source framework for orchestrating AI agents. It enables you to build autonomous, collaborative agent systems that can solve complex tasks through: - **Multi-Agent Orchestration**: Create teams of specialized AI agents that collaborate to solve complex tasks - **Dual Paradigms**: Use both autonomous Crews and controlled Flows for maximum flexibility - **Role-Based Collaboration**: Each agent has a distinct role, goal, and backstory - **Process Management**: Sequential, hierarchical, and custom execution flows - **Tool Integration**: Equip agents with tools to interact with external systems By connecting CrewAI with APIpie, you unlock access to a wide range of powerful language models while leveraging CrewAI's robust agent orchestration capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with CrewAI. ### 4. Install CrewAI Install CrewAI and its tools using pip: ```bash pip install crewai # For additional tools (optional) pip install 'crewai[tools]' ``` ### 5. Configure CrewAI for APIpie Set up your environment to use APIpie with CrewAI: ```python import os from crewai import Agent, Task, Crew, LLM # Set up with environment variables os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Or configure directly in your code llm = LLM( provider="openai", api_key="your-apipie-api-key", base_url="https://apipie.ai/v1", model="gpt-4o-mini" # Or any model available on APIpie ) ``` --- ## Key Features - **Fast Performance**: Lightweight, standalone framework built from scratch for speed - **Dual Paradigms**: - **Crews**: Teams of autonomous agents working together - **Flows**: Precise, event-driven control for complex workflows - **Flexible Customization**: Control both high-level workflows and fine-grained agent behaviors - **Tool Integration**: Extend agents with powerful tools to interact with external services - **Process Management**: Sequential, hierarchical, and custom execution processes --- ## Example Workflows | Application Type | What CrewAI Helps You Build | | --------------------------- | -------------------------------------------------------------- | | Research Workflows | Multi-agent teams that research, analyze, and summarize topics | | Content Creation | Teams that plan, write, and refine documents and creative work | | Data Analysis | Collaborative analysis of financial data and market trends | | Business Process Automation | Complex multi-step workflows with business logic and tools | | Planning Systems | Agents that collaborate on planning trips, events, or projects | --- ## Using CrewAI with APIpie ### Basic Multi-Agent Research Crew ```python import os from crewai import Agent, Task, Crew, LLM # Configure APIpie as the LLM provider os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Alternatively, configure the LLM explicitly llm = LLM( provider="openai", api_key="your-apipie-api-key", base_url="https://apipie.ai/v1", model="gpt-4o-mini" # Or any model available on APIpie ) # Create a researcher agent researcher = Agent( role="Research Analyst", goal="Find and summarize information about specific topics", backstory="You are an experienced researcher with attention to detail", llm=llm, # Use the configured LLM verbose=True ) # Create a writer agent writer = Agent( role="Content Writer", goal="Create engaging and informative content from research", backstory="You are a skilled writer who turns complex information into readable content", llm=llm, # Use the configured LLM verbose=True ) # Create tasks for the agents research_task = Task( description="Conduct thorough research about artificial intelligence trends in 2025", expected_output="A list with 10 bullet points of the most relevant AI trends in 2025", agent=researcher ) writing_task = Task( description="Write a comprehensive article about AI trends in 2025 based on the research provided", expected_output="A 1000-word article with clear sections covering each major trend", agent=writer ) # Create a crew with sequential process crew = Crew( agents=[researcher, writer], tasks=[research_task, writing_task], verbose=True ) # Execute the crew result = crew.kickoff() print(result) ``` ### Adding Tools to Agents ```python import os from crewai import Agent, Task, Crew, LLM from crewai_tools import SerperDevTool # Search tool # Configure your credentials os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" os.environ["SERPER_API_KEY"] = "your-serper-api-key" # For the search tool # Create a researcher agent with search capabilities researcher = Agent( role="Market Researcher", goal="Research market trends and competitor analysis", backstory="You're a skilled market researcher with years of experience in competitor analysis", tools=[SerperDevTool()], # Add the search tool llm=LLM( provider="openai", api_key=os.environ["OPENAI_API_KEY"], base_url=os.environ["OPENAI_API_BASE"], model="gpt-4o" ), verbose=True ) # Create a financial analyst agent analyst = Agent( role="Financial Analyst", goal="Analyze financial data and provide investment recommendations", backstory="You're a financial expert who excels at interpreting market data", llm=LLM( provider="openai", api_key=os.environ["OPENAI_API_KEY"], base_url=os.environ["OPENAI_API_BASE"], model="gpt-4o" ), verbose=True ) # Define tasks research_task = Task( description="Research the current state of AI chip manufacturing. Focus on major players, market share, and recent innovations.", expected_output="A detailed report on the AI chip market landscape with key players and market trends", agent=researcher ) analysis_task = Task( description="Based on the research, analyze the investment potential of top AI chip manufacturers. Consider growth prospects, risks, and competitive advantages.", expected_output="An investment analysis report with recommendations for the top 3 AI chip manufacturers", agent=analyst ) # Create and execute the crew crew = Crew( agents=[researcher, analyst], tasks=[research_task, analysis_task], verbose=True ) result = crew.kickoff() print(result) ``` ### Using CrewAI Flows with APIpie CrewAI Flows provide more precise control over execution paths: ```python import os from crewai import Agent, Task, Crew, LLM from crewai.flow.flow import Flow, listen, start, router from pydantic import BaseModel # Configure APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Define a state model for the flow class ResearchState(BaseModel): topic: str = "" confidence: float = 0.0 findings: list = [] # Create a flow with structured control class ResearchFlow(Flow[ResearchState]): @start() def initialize_research(self): self.state.topic = "Machine Learning Applications in Healthcare" return {"topic": self.state.topic} @listen(initialize_research) def conduct_research(self, inputs): # Create a research agent researcher = Agent( role="Medical Research Specialist", goal="Find and analyze the latest applications of ML in healthcare", backstory="You're a specialist in healthcare technology research", llm=LLM( provider="openai", api_key=os.environ["OPENAI_API_KEY"], base_url=os.environ["OPENAI_API_BASE"], model="gpt-4o-mini" ) ) # Create a research task research_task = Task( description=f"Research the latest applications of machine learning in healthcare, focusing on {inputs['topic']}", expected_output="A comprehensive report on ML applications in healthcare", agent=researcher ) # Create and execute a simple crew research_crew = Crew( agents=[researcher], tasks=[research_task], verbose=True ) result = research_crew.kickoff() # Update state with findings self.state.findings.append(result) self.state.confidence = 0.75 return result @router(conduct_research) def determine_next_step(self): # Conditional routing based on confidence if self.state.confidence > 0.7: return "sufficient_data" else: return "need_more_data" @listen("sufficient_data") def generate_final_report(self): # Create a report writing agent writer = Agent( role="Medical Content Writer", goal="Transform research into readable reports", backstory="You specialize in making complex medical technology understandable", llm=LLM( provider="openai", api_key=os.environ["OPENAI_API_KEY"], base_url=os.environ["OPENAI_API_BASE"], model="gpt-4o-mini" ) ) # Create a report writing task report_task = Task( description="Transform the research findings into a comprehensive report for healthcare professionals", expected_output="A detailed report on ML applications in healthcare with recommendations", agent=writer, context="\n".join(self.state.findings) ) # Create and execute the crew report_crew = Crew( agents=[writer], tasks=[report_task], verbose=True ) return report_crew.kickoff() @listen("need_more_data") def request_additional_research(self): # Logic for when more research is needed return "More research needed to increase confidence in findings" # Execute the flow research_flow = ResearchFlow() final_result = research_flow.start() print(final_result) ``` ### Using CrewAI's YAML Configuration CrewAI also supports defining agents and tasks using YAML files: ```python import os from crewai import Agent, Task, Crew, LLM, Process from crewai.project import CrewBase, agent, crew, task from crewai_tools import SerperDevTool from crewai.agents.agent_builder.base_agent import BaseAgent from typing import List # Configure APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" @CrewBase class ResearchCrew(): """Research crew for analyzing topics""" agents: List[BaseAgent] tasks: List[Task] @agent def researcher(self) -> Agent: return Agent( config=self.agents_config['researcher'], verbose=True, tools=[SerperDevTool()], llm=LLM( provider="openai", api_key=os.environ["OPENAI_API_KEY"], base_url=os.environ["OPENAI_API_BASE"], model="gpt-4o-mini" ) ) @agent def analyst(self) -> Agent: return Agent( config=self.agents_config['analyst'], verbose=True, llm=LLM( provider="openai", api_key=os.environ["OPENAI_API_KEY"], base_url=os.environ["OPENAI_API_BASE"], model="gpt-4o-mini" ) ) @task def research_task(self) -> Task: return Task( config=self.tasks_config['research_task'], ) @task def analysis_task(self) -> Task: return Task( config=self.tasks_config['analysis_task'], output_file='analysis.md' ) @crew def crew(self) -> Crew: """Creates the research crew""" return Crew( agents=self.agents, tasks=self.tasks, process=Process.sequential, verbose=True, ) # Execute the crew with inputs inputs = { 'topic': 'Quantum Computing Applications' } result = ResearchCrew().crew().kickoff(inputs=inputs) print(result) ``` --- ## Troubleshooting & FAQ - **Which models are supported?**:br Any model available via APIpie's OpenAI-compatible endpoint can be used with CrewAI. - **How do I handle environment variables securely?**:br Store your API keys in environment variables or use a secure environment management tool. Never commit API keys to repositories. - **Can I use different models for different agents?**:br Yes, you can create different LLM objects for different agents, allowing you to optimize each agent for its specific task. - **How do I add custom tools to my agents?**:br CrewAI supports custom tools through the `@tool` decorator or by creating a class that inherits from `BaseTool`. - **Can I mix Crews and Flows?**:br Yes, CrewAI is designed to let you use Crews within Flows for maximum flexibility, combining autonomous agent behavior with precise control flow. - **How can I save agent outputs to files?**:br Use the `output_file` parameter when creating a Task to automatically save the agent's output to a file. For more information, see the [CrewAI documentation](https://docs.crewai.com/){rel=""nofollow""} or the [GitHub repository](https://github.com/crewAIInc/crewAI){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # DSPy Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![DSPy](https://apipie.ai/img/docs/integrations/DSPy/dspy_logo.png){height="125"} :: This guide will walk you through integrating DSPy with APIpie, enabling you to build and optimize modular AI systems with Python code instead of hand-crafted prompting. ## What is [DSPy](https://dspy.ai/){rel=""nofollow""}? DSPy (Declarative Self-improving Python) is a framework for programming—not prompting—language models. It enables you to: - **Build Modular AI Systems**: Create compositional Python code instead of brittle prompts - **Optimize Prompts and Weights**: Automatically improve your AI pipelines without manual tuning - **Declarative Programming**: Define what you want your AI to do, not how to do it - **Self-improvement**: DSPy can teach language models to deliver high-quality outputs - **Composable Modules**: Build complex AI systems from reusable components By connecting DSPy with APIpie, you gain access to a wide range of powerful language models while leveraging DSPy's sophisticated programming paradigm and optimization capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with DSPy. ### 4. Install DSPy Install DSPy using pip: ```bash pip install dspy ``` For the latest version from the GitHub repository: ```bash pip install git+https://github.com/stanfordnlp/dspy.git ``` ### 5. Configure DSPy for APIpie DSPy can connect to APIpie through the OpenAI-compatible interface: ```python import os import dspy # Set environment variables for APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Configure DSPy to use APIpie lm = dspy.OpenAI(model="gpt-4o-mini") # Use any model available on APIpie dspy.configure(lm=lm) ``` Alternatively, you can configure the language model explicitly: ```python import dspy # Configure DSPy with explicit parameters lm = dspy.OpenAI( model="gpt-4o-mini", # Use any model available on APIpie api_key="your-apipie-api-key", api_base="https://apipie.ai/v1" ) # Set this as the default language model dspy.configure(lm=lm) ``` --- ## Key Features - **Declarative Programming**: Define what you want, not how to get it - **Modular Components**: Build complex systems from reusable modules - **Automatic Optimization**: Improve prompts and weights automatically - **Compositional Design**: Combine modules to create sophisticated pipelines - **Tool Integration**: Seamlessly integrate external tools and knowledge sources - **RAG Integration**: Native support for retrieval-augmented generation --- ## Example Workflows | Application Type | What DSPy Helps You Build | | -------------------------- | -------------------------------------------------------------- | | Question Answering Systems | Building QA systems that reason over text and data | | Multi-step Reasoning | Creating step-by-step reasoning pipelines for complex problems | | Information Retrieval | Optimizing retrieval-augmented generation systems | | Text Classification | Developing robust classifiers with automatic prompt tuning | | Agent Frameworks | Building agents that can use tools and improve over time | --- ## Using DSPy with APIpie ### Basic Question Answering ```python import os import dspy # Configure DSPy with APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Set up the language model lm = dspy.OpenAI(model="gpt-4o-mini") dspy.configure(lm=lm) # Define a simple question-answering module class BasicQA(dspy.Module): def __init__(self): super().__init__() self.generate_answer = dspy.ChainOfThought("question -> answer") def forward(self, question): return self.generate_answer(question=question) # Create and use the QA module qa_module = BasicQA() response = qa_module("What is the capital of France?") print(response.answer) # Paris ``` ### Using DSPy with External Tools ```python import os import dspy from typing import List # Configure DSPy with APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Set up the language model lm = dspy.OpenAI(model="gpt-4o") dspy.configure(lm=lm) # Define a tool for math operations def evaluate_math(expression: str) -> float: """Evaluate a mathematical expression.""" return float(eval(expression)) # Define a simple information retrieval tool def search_web(query: str) -> List[str]: """Search the web for information.""" # In a real application, this would use a search API # This is just a mock example if "Paris" in query: return ["Paris is the capital of France.", "Paris has a population of about 2.2 million."] elif "Rome" in query: return ["Rome is the capital of Italy.", "Rome was founded in 753 BC."] else: return ["No specific information found."] # Create a ReAct agent that can use these tools react_agent = dspy.ReAct( "question -> answer", tools=[evaluate_math, search_web] ) # Use the agent to answer questions result1 = react_agent(question="What is 123 * 456?") print(f"Math answer: {result1.answer}") result2 = react_agent(question="What is the capital of Italy?") print(f"Information answer: {result2.answer}") ``` ### Building and Optimizing a RAG System ```python import os import dspy from dspy.retrieve import ColBERTv2 # Configure DSPy with APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Set up the language model lm = dspy.OpenAI(model="gpt-4o-mini") dspy.configure(lm=lm) # Mock retriever for demonstration # In a real application, you would use a proper retrieval system # For example: retriever = ColBERTv2(url="http://your-colbert-server") class MockRetriever: def __call__(self, query, k=3): # Simulate document retrieval if "climate change" in query.lower(): docs = [ {"text": "Climate change is a significant global challenge that requires immediate action."}, {"text": "Rising temperatures are leading to more extreme weather events worldwide."}, {"text": "Renewable energy is a key solution to address climate change."} ] else: docs = [ {"text": "No specific information found for this query."} ] return docs[:k] retriever = MockRetriever() # Define a RAG module using DSPy class RAG(dspy.Module): def __init__(self, retriever): super().__init__() self.retriever = retriever self.generate_query = dspy.Predict("question -> query") self.generate_answer = dspy.ChainOfThought("question, context -> answer") def forward(self, question): # Generate an effective query query = self.generate_query(question=question).query # Retrieve relevant documents docs = self.retriever(query) context = "\n".join([d["text"] for d in docs]) # Generate an answer based on the retrieved information return self.generate_answer(question=question, context=context) # Create the RAG system rag_system = RAG(retriever) # Optimize the RAG system teleprompter = dspy.Teleprompter(introspect=True) optimized_rag = teleprompter.optimize( rag_system, task_demos=[ dspy.Example( question="What are the main causes of climate change?", answer="The main causes of climate change include greenhouse gas emissions from burning fossil fuels, deforestation, and industrial processes." ), dspy.Example( question="How can we address climate change?", answer="Climate change can be addressed through reducing greenhouse gas emissions, transitioning to renewable energy, improving energy efficiency, and implementing sustainable practices." ) ] ) # Use the optimized RAG system result = optimized_rag("What are the effects of climate change on agriculture?") print(result.answer) ``` ### Advanced Multi-Step Reasoning ```python import os import dspy # Configure DSPy with APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Set up the language model lm = dspy.OpenAI(model="gpt-4o") dspy.configure(lm=lm) # Define signatures for multi-step reasoning class GeneratePlan(dspy.Signature): """Generate a step-by-step plan to solve a complex problem.""" problem = dspy.InputField() plan = dspy.OutputField(desc="A detailed step-by-step plan with 3-5 steps") class ExecuteStep(dspy.Signature): """Execute a specific step in the problem-solving plan.""" problem = dspy.InputField() previous_steps = dspy.InputField(desc="Results from previous steps, if any") current_step = dspy.InputField(desc="The current step to execute") result = dspy.OutputField(desc="The result of executing the current step") class ProvideFinalAnswer(dspy.Signature): """Provide the final answer based on all steps executed.""" problem = dspy.InputField() all_steps = dspy.InputField(desc="All steps and their results") answer = dspy.OutputField(desc="The final answer to the problem") # Create a multi-step reasoning module class MultistepReasoning(dspy.Module): def __init__(self): super().__init__() self.generate_plan = dspy.ChainOfThought(GeneratePlan) self.execute_step = dspy.ChainOfThought(ExecuteStep) self.provide_final_answer = dspy.ChainOfThought(ProvideFinalAnswer) def forward(self, problem): # Generate a plan plan_output = self.generate_plan(problem=problem) # Parse the plan into steps steps = [step.strip() for step in plan_output.plan.split('\n') if step.strip()] # Execute each step previous_results = [] all_results = [] for i, step in enumerate(steps): # Execute the current step step_result = self.execute_step( problem=problem, previous_steps='\n'.join(previous_results), current_step=step ) # Store the result result_text = f"Step {i+1}: {step}\nResult: {step_result.result}" previous_results.append(result_text) all_results.append(result_text) # Provide the final answer final_answer = self.provide_final_answer( problem=problem, all_steps='\n'.join(all_results) ) return final_answer.answer # Use the multi-step reasoning module reasoning_module = MultistepReasoning() answer = reasoning_module("Calculate the approximate area of a circle with diameter 10 cm.") print(answer) ``` --- ## Troubleshooting & FAQ - **Which models are supported?**:br Any model available via APIpie's OpenAI-compatible endpoint can be used with DSPy. - **How do I handle environment variables securely?**:br Store your API keys in environment variables or use a secure environment management tool. Never commit API keys to repositories. - **Can I use DSPy's optimization capabilities with APIpie?**:br Yes, DSPy's Teleprompter and other optimization techniques work seamlessly with APIpie's models. - **How do I integrate external tools with DSPy?**:br DSPy supports tool integration through its ReAct module or by defining custom Python functions that DSPy modules can call. - **Can I use DSPy with retrieval systems?**:br Yes, DSPy has native support for retrieval-augmented generation (RAG) and can work with vector databases and retrieval systems. - **How do I save optimized prompts?**:br After optimizing prompts with DSPy's Teleprompter, you can extract and save the optimized prompts for deployment. For more information, see the [DSPy documentation](https://dspy.ai/){rel=""nofollow""} or the [GitHub repository](https://github.com/stanfordnlp/dspy){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Google Agent Development Kit (ADK) Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![Google Agent Development Kit](https://apipie.ai/img/docs/integrations/Google/agent-development-kit.png){height="125"} :: This guide will walk you through integrating Google's Agent Development Kit (ADK) with APIpie, enabling you to build sophisticated AI agents with flexible tooling and deployment options. ## What is [Google Agent Development Kit](https://google.github.io/adk-docs/){rel=""nofollow""}? Google Agent Development Kit (ADK) is an open-source, code-first toolkit for building, evaluating, and deploying AI agents. It provides: - **Rich Tool Ecosystem**: Utilize pre-built tools, custom functions, OpenAPI specs, or integrate existing tools - **Code-First Development**: Define agent logic, tools, and orchestration directly in Python - **Modular Multi-Agent Systems**: Design scalable applications with specialized agents in flexible hierarchies - **Deployment Options**: Deploy agents on Cloud Run or Vertex AI Agent Engine - **Model Agnosticism**: Works with Gemini models by default, but also supports other LLMs By connecting ADK to APIpie, you gain access to a wide range of powerful language models while leveraging ADK's sophisticated agent development capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with Google ADK. ### 4. Install Google Agent Development Kit ```bash pip install google-adk ``` For additional capabilities (optional): ```bash # For tracing and monitoring pip install google-adk[tracing] # For development UI pip install google-adk[web] # For voice support pip install google-adk[voice] ``` ### 5. Configure ADK for APIpie ADK is designed to work with Google's Gemini models by default, but it can be configured to use alternative LLM providers like APIpie. You can set this up by creating a custom LLM provider class: ```python import os from google.adk.agents import Agent from google.adk.llms import LLM from google.adk.llms.openai import OpenAILLM # Configure a custom LLM with APIpie apipie_llm = OpenAILLM( model="gpt-4o", # Use any model available on APIpie api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1" ) # Create an agent with the custom LLM root_agent = Agent( name="apipie_agent", llm=apipie_llm, description="Agent that uses APIpie as the LLM provider", instruction="You are a helpful assistant that provides accurate and concise information." ) ``` --- ## Key Features - **Tool Integration**: Easily add custom tools, functions, and APIs to your agents - **Code-First Approach**: Build agents directly in Python with full software engineering practices - **Multi-Agent Architecture**: Create specialized agents that work together on complex tasks - **Streaming Support**: Real-time streaming for text, voice, and video interactions - **Flexible Deployment**: Run agents locally, in containers, or on cloud platforms - **Provider Agnosticism**: Works with any LLM through custom providers --- ## Example Workflows | Application Type | What Google ADK Helps You Build | | ------------------------ | ---------------------------------------------------------------- | | Customer Support Systems | Agents that can access knowledge bases and respond to queries | | Research Assistants | Multi-tool agents that search, summarize, and analyze data | | Voice/Video Applications | Interactive applications using streaming for real-time responses | | Enterprise Workflows | Complex business processes with specialized agent teams | | Data Analysis Agents | Systems that process, visualize, and interpret structured data | --- ## Using Google ADK with APIpie ### Basic Agent with Multiple Tools ```python import os import datetime from zoneinfo import ZoneInfo from google.adk.agents import Agent from google.adk.llms.openai import OpenAILLM # Configure APIpie as the LLM provider apipie_llm = OpenAILLM( model="gpt-4o", api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1" ) def get_weather(city: str) -> dict: """Retrieves the current weather report for a specified city. Args: city (str): The name of the city for which to retrieve the weather report. Returns: dict: status and result or error msg. """ if city.lower() == "new york": return { "status": "success", "report": ( "The weather in New York is sunny with a temperature of 25 degrees" " Celsius (77 degrees Fahrenheit)." ), } else: return { "status": "error", "error_message": f"Weather information for '{city}' is not available.", } def get_current_time(city: str) -> dict: """Returns the current time in a specified city. Args: city (str): The name of the city for which to retrieve the current time. Returns: dict: status and result or error msg. """ if city.lower() == "new york": tz_identifier = "America/New_York" else: return { "status": "error", "error_message": f"Sorry, I don't have timezone information for {city}.", } tz = ZoneInfo(tz_identifier) now = datetime.datetime.now(tz) report = f'The current time in {city} is {now.strftime("%Y-%m-%d %H:%M:%S %Z%z")}' return {"status": "success", "report": report} # Create agent with multiple tools root_agent = Agent( name="weather_time_agent", llm=apipie_llm, description="Agent to answer questions about the time and weather in a city.", instruction="You are a helpful agent who can answer user questions about the time and weather in a city.", tools=[get_weather, get_current_time], ) ``` ### Using Built-in Google Search Tool ```python import os from google.adk.agents import Agent from google.adk.tools import google_search from google.adk.llms.openai import OpenAILLM # Configure APIpie as the LLM provider apipie_llm = OpenAILLM( model="gpt-4o", api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1" ) # Create a search agent search_agent = Agent( name="search_agent", llm=apipie_llm, description="Agent to answer questions using Google Search.", instruction="You are an expert researcher. You always stick to the facts and cite your sources.", tools=[google_search], ) ``` ### Multi-Agent System with Specialized Agents ```python import os from google.adk.agents import Agent from google.adk.agents.sequential_agent import SequentialAgent from google.adk.llms.openai import OpenAILLM # Configure APIpie as the LLM provider apipie_llm = OpenAILLM( model="gpt-4o", api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1" ) # Create specialized agents research_agent = Agent( name="research_agent", llm=apipie_llm, description="Agent to conduct research on topics.", instruction="You are a research specialist who finds accurate information." ) analysis_agent = Agent( name="analysis_agent", llm=apipie_llm, description="Agent to analyze research findings.", instruction="You are an analysis expert who synthesizes information into insights." ) presentation_agent = Agent( name="presentation_agent", llm=apipie_llm, description="Agent to present analysis in a clear format.", instruction="You are a communication expert who presents complex information clearly." ) # Create a sequential agent that chains these specialized agents agent_workflow = SequentialAgent( name="research_workflow", llm=apipie_llm, description="Research workflow that researches, analyzes, and presents information.", instruction="Coordinate research, analysis, and presentation to provide comprehensive answers.", agents=[research_agent, analysis_agent, presentation_agent] ) ``` ### Streaming Support for Voice and Video ```python import os from google.adk.agents import Agent, LiveRequestQueue from google.adk.runners import Runner from google.adk.agents.run_config import RunConfig from google.adk.sessions.in_memory_session_service import InMemorySessionService from google.adk.llms.openai import OpenAILLM # Configure APIpie as the LLM provider apipie_llm = OpenAILLM( model="gpt-4o", # Use a model that supports streaming api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1" ) # Create a streaming-capable agent streaming_agent = Agent( name="streaming_agent", llm=apipie_llm, description="Agent that supports real-time streaming for voice and video.", instruction="You respond naturally and helpfully to voice and video inputs." ) # Set up session and runner for streaming session_service = InMemorySessionService() session = session_service.create_session( app_name="Streaming Demo", user_id="user_123", session_id="session_456" ) # Create a runner runner = Runner( app_name="Streaming Demo", agent=streaming_agent, session_service=session_service ) # Configure for audio response run_config = RunConfig(response_modalities=["TEXT", "AUDIO"]) # Create a live request queue for two-way communication live_request_queue = LiveRequestQueue() # Start a streaming session live_events = runner.run_live( session=session, live_request_queue=live_request_queue, run_config=run_config ) # In an async context, you would process live_events and live_request_queue # to handle the streaming communication ``` --- ## Running Your Agents ADK provides multiple ways to interact with your agents: ### Using the Development UI ```bash # Navigate to your agent's parent directory cd path/to/your/project # Launch the dev UI adk web ``` This will start a web interface where you can interact with your agent. ### Using the Terminal ```bash # Navigate to your agent's parent directory cd path/to/your/project # Run the agent in terminal mode adk run your_agent_module ``` ### As an API Server ```bash # Navigate to your agent's parent directory cd path/to/your/project # Start the API server adk api_server your_agent_module ``` This will start an API server that you can integrate with web applications or other services. --- ## Troubleshooting & FAQ - **Can I use APIpie models with Google ADK?**:br Yes, ADK is designed to be model-agnostic. You can use the OpenAILLM provider to connect to any OpenAI-compatible API like APIpie. - **How do I handle environment variables securely?**:br Store your API keys in a `.env` file and load them using python-dotenv. Never commit API keys to your repositories. - **What if I need to use multiple models in the same agent system?**:br You can create different LLM instances and assign them to different agents in your system based on their requirements. - **How do I debug agent behavior?**:br Use the `adk web` development UI to inspect agent interactions, tool calls, and responses. Enable tracing for more detailed insights. - **Can I deploy my agent to a production environment?**:br Yes, ADK supports containerization for deployment on platforms like Cloud Run or Vertex AI Agent Engine. - **Are there any limitations when using non-Google LLMs?**:br Some ADK features may be optimized for Gemini models, but core functionality works with any LLM. Check the compatibility notes for specific features. For more information, see the [Google ADK documentation](https://google.github.io/adk-docs/){rel=""nofollow""} or the [GitHub repository](https://github.com/google/adk-docs){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Hugging Face Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![Hugging Face](https://apipie.ai/img/docs/integrations/Huggingface/hf-logo.svg){height="125"} :: This guide will walk you through integrating Hugging Face with APIpie, enabling you to use thousands of open-source AI models for various tasks through a unified interface. ## What is [Hugging Face](https://huggingface.co/){rel=""nofollow""}? Hugging Face is a leading platform for the AI community, providing: - **Open-Source Models**: Access to thousands of pre-trained models for various tasks - **Inference API**: A unified interface to run inference across multiple models - **Model Hub**: A repository of models, datasets, and spaces for the community - **Transformers**: A popular library for natural language processing (NLP) tasks - **Datasets**: A collection of ready-to-use datasets for AI training - **Spaces**: Interactive demos of AI applications By connecting Hugging Face with APIpie, you can access a wide range of models through a consistent API, while benefiting from APIpie's additional features like model routing and failover. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with Hugging Face. ### 4. Install Required Libraries For Python: ```bash pip install huggingface_hub openai ``` For JavaScript/TypeScript: ```bash npm install @huggingface/inference openai # or yarn add @huggingface/inference openai # or pnpm add @huggingface/inference openai ``` ### 5. Configure Hugging Face with APIpie There are two main ways to integrate Hugging Face with APIpie: #### A. Using Hugging Face's InferenceClient with APIpie as a Provider ```python from huggingface_hub import InferenceClient # Initialize the client with APIpie as the provider client = InferenceClient( provider="openai", api_key="your-apipie-api-key", base_url="https://apipie.ai/v1", ) ``` #### B. Using OpenAI SDK with APIpie (Recommended) ```python import os import openai # Configure OpenAI to use APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Create a client with APIpie configuration client = openai.OpenAI() ``` --- ## Key Features - **Thousands of Models**: Access to a vast range of pre-trained models - **Multi-Modal Support**: Models for text, images, audio, and more - **Unified API**: Consistent interface for different types of models - **Open Source**: Many models are open-source and can be deployed anywhere - **Community Driven**: Access to models developed by the AI community - **Inference Options**: Run inference in the cloud or deploy locally --- ## Example Workflows | Application Type | What Hugging Face Helps You Build | | --------------------------- | ---------------------------------------------------------- | | Natural Language Processing | Text generation, translation, summarization, and more | | Computer Vision | Image classification, object detection, image segmentation | | Audio Processing | Speech recognition, audio classification, text-to-speech | | Multimodal Applications | Vision-language tasks, document understanding | | Specialized Models | Domain-specific models like biomedical or financial NLP | --- ## Using Hugging Face with APIpie ### Text Generation with Chat Completion ```python from huggingface_hub import InferenceClient # Initialize client with APIpie client = InferenceClient( provider="openai", api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) # Use chat completion API with Hugging Face models through APIpie response = client.chat.completions.create( model="meta-llama/Meta-Llama-3-8B-Instruct", # Hugging Face model messages=[ {"role": "user", "content": "What is the capital of France?"} ], max_tokens=100 ) print(response.choices[0].message.content) ``` ### JavaScript Example with Chat Completion ```javascript import { HfInference } from '@huggingface/inference'; // Initialize with APIpie configuration const hf = new HfInference({ apiKey: 'your-apipie-api-key', baseURL: 'https://apipie.ai/v1', }); // Chat completion with Hugging Face model const chatCompletion = await hf.chatCompletion({ model: 'meta-llama/Meta-Llama-3-8B-Instruct', messages: [ { role: 'user', content: 'What is the capital of France?', }, ], max_tokens: 100, }); console.log(chatCompletion.choices[0].message.content); ``` ### Text-to-Image Generation ```python import openai import os from PIL import Image import io import base64 # Configure OpenAI client to use APIpie client = openai.OpenAI( api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) # Generate an image using a Hugging Face model through APIpie response = client.images.generate( model="stabilityai/stable-diffusion-2-1", # Hugging Face model prompt="A serene landscape with mountains and a lake at sunset", n=1, size="1024x1024" ) # Process and display the image image_url = response.data[0].url # Or if you get a base64 encoded image # image_data = base64.b64decode(response.data[0].b64_json) # image = Image.open(io.BytesIO(image_data)) # image.save("generated_image.png") ``` ### Audio Transcription ```python import openai import os # Configure client to use APIpie client = openai.OpenAI( api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) # Transcribe audio using a Hugging Face model with open("audio_sample.mp3", "rb") as audio_file: response = client.audio.transcriptions.create( model="facebook/wav2vec2-large-960h-lv60-self", # Hugging Face model file=audio_file, response_format="text" ) print(response) ``` ### Embeddings Generation ```python from huggingface_hub import InferenceClient # Initialize client with APIpie client = InferenceClient( provider="openai", api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) # Generate embeddings response = client.embeddings.create( model="sentence-transformers/all-MiniLM-L6-v2", # Hugging Face model input="The food was delicious and the service was excellent." ) print(response.data[0].embedding) ``` ### Multi-Modal Tasks ```python import openai import os import base64 from PIL import Image import io # Configure client to use APIpie client = openai.OpenAI( api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) # Read image file with open("image.jpg", "rb") as image_file: encoded_image = base64.b64encode(image_file.read()).decode('utf-8') # Perform visual question answering using a Hugging Face model response = client.chat.completions.create( model="llava-hf/llava-1.5-7b-hf", # Hugging Face model messages=[ { "role": "user", "content": [ {"type": "text", "text": "What is shown in this image?"}, {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{encoded_image}"}} ] } ], max_tokens=300 ) print(response.choices[0].message.content) ``` --- ## Troubleshooting & FAQ - **Which Hugging Face models are supported?**:br APIpie supports a wide range of Hugging Face models. The specific models available depend on your APIpie subscription tier. - **How do I handle environment variables securely?**:br Store your API keys in environment variables or use a secure environment management tool. Never commit API keys to repositories. - **Can I use my own Hugging Face models with APIpie?**:br Yes, if your model is hosted on the Hugging Face Hub and is supported by APIpie, you can use it through the integration. For private models, you may need to provide additional authentication. - **How do I handle rate limits?**:br APIpie has its own rate limits based on your subscription tier. Check the APIpie documentation for specific limits and how to handle them. - **What if a Hugging Face model is not available through APIpie?**:br You can request support for specific models through APIpie's support channels. Alternatively, you can use the Hugging Face Inference API directly for models not available through APIpie. - **Is there a difference in latency when using Hugging Face through APIpie?**:br There might be a minimal latency overhead when using APIpie as an intermediary, but this is often offset by APIpie's routing and caching capabilities, especially for frequently used requests. For more information, see the [Python Inference client Documentation](https://huggingface.co/docs/huggingface_hub/main/en/package_reference/inference_client){rel=""nofollow""} or the [Java Script Inference client Documentation](https://huggingface.co/docs/huggingface.js/inference/README){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Agno Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![Agno](https://apipie.ai/img/docs/integrations/Agno/logo-dark.svg){height="125"} :: This guide will walk you through integrating Agno with APIpie, enabling you to build powerful AI agents with advanced reasoning, multimodal capabilities, and tools integration. ## What is [Agno](https://docs.agno.com){rel=""nofollow""}? Agno is a lightweight library for building intelligent AI agents with memory, knowledge, tools, and reasoning. It provides a flexible and powerful system for: - **Multimodal agents** that can process text, images, audio, and video - **Reasoning capabilities** to help agents "think" through complex problems - **Advanced tools integration** with 20+ pre-built tools - **Team-based agent architecture** with routing and coordination - **Knowledge integration** via vector databases for RAG applications - **Long-term memory** and session storage By connecting Agno with APIpie, you unlock access to a wide range of powerful language models while leveraging Agno's robust agent architecture and tooling. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with Agno. ### 4. Install Agno Install Agno using pip: ```bash pip install -U agno ``` For specific tools, you may need to install additional packages. For example: ```bash # For web search capabilities pip install -U agno duckduckgo-search # For finance tools pip install -U agno yfinance # For PDF processing pip install -U agno pypdf ``` ### 5. Configure Agno for APIpie Agno can use any API-based language model, including those from APIpie. Here's how to configure it: ```python from agno.agent import Agent from agno.models.openai import OpenAIChat # Create a custom model configuration for APIpie class APIpieModel(OpenAIChat): def __init__(self, id: str = "gpt-4o-mini", **kwargs): super().__init__( id=id, api_key="your-apipie-api-key", base_url="https://apipie.ai/v1", **kwargs ) # Create an agent with APIpie model agent = Agent( model=APIpieModel(), markdown=True ) ``` You can also use environment variables: ```python import os from agno.agent import Agent from agno.models.openai import OpenAIChat # Set environment variables for APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Create an agent with the default configuration agent = Agent( model=OpenAIChat(id="gpt-4o-mini"), markdown=True ) ``` --- ## Key Features - **Lightning Fast Performance**: Agno agents instantiate in microseconds and use minimal memory - **Model Agnostic**: Connect to 23+ model providers, including APIpie - **Advanced Reasoning**: Make agents "think" through problems with specialized reasoning tools - **Multimodal Capabilities**: Process and generate text, images, audio, and video - **Agent Teams**: Build multi-agent systems with specialized roles and coordination - **Knowledge Integration**: Connect to vector databases for Agentic RAG applications - **Pre-Built API Routes**: Serve agents via FastAPI with minimal setup --- ## Example Workflows | Application Type | What Agno Helps You Build | | ------------------------- | ------------------------------------------------------------- | | Reasoning Agents | Agents that can think through complex problems step-by-step | | Multimodal Assistants | Agents that process text, images, audio, and video | | Research & Analysis Tools | Agents that gather, analyze, and synthesize information | | Financial Analysis | Agents that analyze stocks, financial data, and market trends | | Multi-Agent Teams | Specialized agent groups that collaborate on complex tasks | --- ## Using Agno with APIpie ### Basic Agent Setup ```python import os from agno.agent import Agent from agno.models.openai import OpenAIChat # Configure APIpie as the model provider os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Create a basic agent agent = Agent( model=OpenAIChat(id="gpt-4o-mini"), description="You are a helpful AI assistant specializing in answering general knowledge questions.", markdown=True ) # Use the agent agent.print_response("What are the main factors contributing to climate change?", stream=True) ``` ### Agent with Web Search Tools ```python import os from agno.agent import Agent from agno.models.openai import OpenAIChat from agno.tools.duckduckgo import DuckDuckGoTools # Configure APIpie as the model provider os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Create an agent with web search capability agent = Agent( model=OpenAIChat(id="gpt-4o-mini"), description="You are an enthusiastic news reporter with a flair for storytelling!", tools=[DuckDuckGoTools()], show_tool_calls=True, markdown=True ) # Get the latest news agent.print_response("Tell me about the latest AI research breakthroughs.", stream=True) ``` ### Reasoning Agent with Financial Analysis ```python import os from agno.agent import Agent from agno.models.openai import OpenAIChat from agno.tools.reasoning import ReasoningTools from agno.tools.yfinance import YFinanceTools # Configure APIpie as the model provider os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Create a reasoning agent for financial analysis agent = Agent( model=OpenAIChat(id="gpt-4o-mini"), tools=[ ReasoningTools(add_instructions=True), YFinanceTools( stock_price=True, analyst_recommendations=True, company_info=True, company_news=True, ), ], instructions=[ "Use tables to display data", "Only output the report, no other text", ], markdown=True, ) # Generate a financial report agent.print_response( "Write a report on NVDA", stream=True, show_full_reasoning=True, stream_intermediate_steps=True, ) ``` ### Agent with Knowledge Base (RAG) ```python import os from agno.agent import Agent from agno.models.openai import OpenAIChat from agno.embedder.openai import OpenAIEmbedder from agno.tools.duckduckgo import DuckDuckGoTools from agno.knowledge.pdf_url import PDFUrlKnowledgeBase from agno.vectordb.lancedb import LanceDb, SearchType # Configure APIpie as the model provider os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Configure embeddings with APIpie os.environ["OPENAI_EMBEDDINGS_API_KEY"] = "your-apipie-api-key" # Same or different key os.environ["OPENAI_EMBEDDINGS_API_BASE"] = "https://apipie.ai/v1" # Create an agent with knowledge base agent = Agent( model=OpenAIChat(id="gpt-4o-mini"), description="You are a medical knowledge expert!", instructions=[ "Search your knowledge base for medical information.", "If the question is better suited for the web, search the web to fill in gaps.", "Prefer the information in your knowledge base over the web results." ], knowledge=PDFUrlKnowledgeBase( urls=["https://example.com/path/to/medical-handbook.pdf"], vector_db=LanceDb( uri="tmp/lancedb", table_name="medical", search_type=SearchType.hybrid, embedder=OpenAIEmbedder(id="text-embedding-3-small"), ), ), tools=[DuckDuckGoTools()], show_tool_calls=True, markdown=True ) # Load the knowledge base (only needed once) if agent.knowledge is not None: agent.knowledge.load() # Query the agent agent.print_response("What are the symptoms of type 2 diabetes?", stream=True) ``` ### Multi-Agent Team ```python import os from agno.agent import Agent from agno.models.openai import OpenAIChat from agno.tools.duckduckgo import DuckDuckGoTools from agno.tools.yfinance import YFinanceTools from agno.team import Team # Configure APIpie as the model provider os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Create specialized agents web_agent = Agent( name="Web Agent", role="Search the web for information", model=OpenAIChat(id="gpt-4o-mini"), tools=[DuckDuckGoTools()], instructions="Always include sources", show_tool_calls=True, markdown=True, ) finance_agent = Agent( name="Finance Agent", role="Get financial data", model=OpenAIChat(id="gpt-4o-mini"), tools=[YFinanceTools(stock_price=True, analyst_recommendations=True, company_info=True)], instructions="Use tables to display data", show_tool_calls=True, markdown=True, ) # Create a team of agents agent_team = Team( mode="coordinate", # Options: "route", "collaborate", "coordinate" members=[web_agent, finance_agent], model=OpenAIChat(id="gpt-4o-mini"), success_criteria="A comprehensive financial news report with clear sections and data-driven insights.", instructions=["Always include sources", "Use tables to display data"], show_tool_calls=True, markdown=True, ) # Run the team agent_team.print_response("What's the market outlook and performance of renewable energy companies?", stream=True) ``` --- ## Serving Agents via API Agno provides built-in FastAPI integration to serve your agents via API: ```python from fastapi import FastAPI from agno.agent import Agent from agno.models.openai import OpenAIChat from agno.api.routes import register_agent_routes # Configure APIpie as the model provider import os os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Create your agent agent = Agent( model=OpenAIChat(id="gpt-4o-mini"), description="You are a helpful AI assistant.", markdown=True ) # Create FastAPI app app = FastAPI(title="My Agno API") # Register agent routes register_agent_routes(app, agent) # Run with: uvicorn app:app --reload ``` --- ## Troubleshooting & FAQ - **Which models are supported?**:br Any model available via APIpie's OpenAI-compatible endpoint can be used with Agno. - **How do I handle environment variables securely?**:br Store your API keys in environment variables or use a secure environment management solution. Never commit API keys to repositories. - **Can I use different models for different agents in a team?**:br Yes, each agent in a team can use a different model, even from different providers. - **How do I monitor my agents?**:br Agno provides monitoring capabilities through [app.agno.com](https://app.agno.com){rel=""nofollow""} if you register your agents. - **Can I use APIpie's routing capabilities with Agno?**:br Yes, APIpie's routing can help you access different models while maintaining the same endpoint structure. This is particularly useful for cost optimization and failover scenarios. For more information, see the [Agno documentation](https://docs.agno.com/){rel=""nofollow""} or the [GitHub repository](https://github.com/agno-agi/agno){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # LangChain Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![LangChain](https://apipie.ai/img/docs/integrations/Langchain/langchain.svg){width="600"} :: This guide will walk you through integrating LangChain with APIpie, enabling you to build powerful AI applications using language models with a flexible and robust framework. ## What is LangChain? LangChain is a popular framework for developing applications powered by language models. It enables applications that: - Connect language models to various data sources and knowledge bases - Create chains of operations for complex reasoning and text processing - Build agents that can make decisions and take actions - Develop sophisticated retrieval-augmented generation (RAG) applications By connecting LangChain to APIpie, you unlock access to a wide range of powerful models, enhanced context windows, and cost-efficient AI capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with LangChain. ### 4. Install LangChain **For JavaScript/TypeScript:** ```TypeScript npm install -S langchain # or yarn add langchain # or pnpm add langchain ``` **For Python:** ```Python pip install -U langchain pip install python-dotenv # Optional, for .env file support ``` ### 5. Configure LangChain for APIpie LangChain is designed to work with OpenAI-compatible endpoints like APIpie. You just need to set the right base URL and API key: **For JavaScript/TypeScript:** ```TypeScript export OPENAI_API_KEY="your-APIpie-key-here" export OPENAI_API_BASE_URL="https://apipie.ai/v1" ``` **For Python:** ```python export OPENAI_API_KEY="your-APIpie-key-here" export OPENAI_API_BASE="https://apipie.ai/v1" ``` ::tip Add these environment variables to your shell profile or a .env file for persistence. :: --- ## Key Features - **Modular Framework:** LangChain's modular design makes it easy to swap components as needed. - **Rich Ecosystem:** Access a wide variety of tools, retrievers, and memory components. - **Model Agnostic:** Work with any language model through consistent interfaces. - **Robust Abstractions:** Built-in patterns for common AI application architectures. - **Model Flexibility:** Access APIpie's latest tools capable models. - **Cost Efficiency:** Control your spending with APIpie's transparent pricing. --- ## Example Workflows | Application Type | What LangChain Helps You Build | | ------------------ | ----------------------------------------------------------------- | | Question Answering | Systems that answer questions using specific documents or data | | Chatbots | Interactive conversational agents with memory and context | | Data Analysis | Analyze and extract insights from structured or unstructured data | | Content Generation | Create articles, summaries, marketing copy with specific styles | | Function Calling | Use LLMs to determine when and how to call external functions | | Agents | Autonomous systems that can reason and take actions on their own | --- ## Using LangChain with APIpie ### JavaScript/TypeScript ```typescript const chat = new ChatOpenAI( { modelName: '', // e.g., 'gpt-4o-mini temperature: 0.8, streaming: true, openAIApiKey: '${APIPIE_API_KEY}', configuration: { baseURL: 'https://apipie.ai/v1', }, }, { basePath: 'https://apipie.ai/v1', }, ); ``` ### Python ```python from langchain.chat_models import ChatOpenAI from langchain.prompts import PromptTemplate from langchain.chains import LLMChain from os import getenv from dotenv import load_dotenv load_dotenv() template = """Question: {question} Answer: Answer like a science teacher.""" prompt = PromptTemplate(template=template, input_variables=["question"]) llm = ChatOpenAI( openai_api_key=getenv("APIPIE_API_KEY"), openai_api_base=getenv("APIPIE_BASE_URL"), model_name="", ) llm_chain = LLMChain(prompt=prompt, llm=llm) question = "Why is the sky blue?" print(llm_chain.run(question)) ``` --- ## Building Interactive AI Apps with Streamlit and LangChain [Streamlit](https://streamlit.io/){rel=""nofollow""} is a powerful Python library that makes it easy to create beautiful, interactive web applications for machine learning and data science. Combined with LangChain and APIpie, you can quickly build sophisticated AI applications with minimal code. ### Why Use Streamlit with LangChain? - **Rapid Development**: Build full-featured AI applications in hours, not weeks - **Interactive UI**: Create rich, responsive interfaces without frontend expertise - **Deployment Simplicity**: Easy deployment options including Streamlit Cloud - **Real-Time Feedback**: Get immediate visual feedback as you develop - **Component Ecosystem**: Access a wide variety of pre-built components ### Simple Streamlit Chat Interface ```python import streamlit as st from langchain.chat_models import ChatOpenAI from components.Sidebar import sidebar from shared import constants from langchain.schema import ( HumanMessage, ) st.title("Langchain Streamlit App") # Add a sidebar for configuration api_key, selected_model = sidebar(constants.APIPIE_DEFAULT_CHAT_MODEL) def generate_response(input_text): chat = ChatOpenAI( temperature=0.7, model=selected_model, openai_api_key=api_key, openai_api_base=constants.APIPIE_API_BASE, ) resp = chat([HumanMessage(content=input_text)]) st.write(resp.content) with st.form("test_form"): text = st.text_area( "Input Text:", "Why is the sky blue?" ) submitted = st.form_submit_button("Enter") if submitted: with st.spinner("Generating response..."): generate_response(text) ``` ### Advanced Streamlit Features For more sophisticated applications, you can enhance your Streamlit app with: - **Chat History**: Store and display conversation history - **File Uploading**: Process documents with `st.file_uploader` - **Interactive Visualizations**: Display charts and graphs of LLM analysis - **Session State**: Maintain state between reruns with `st.session_state` - **Multipage Apps**: Create applications with multiple pages ### Deploying Your Streamlit App Once you've built your LangChain + Streamlit application: 1. **Local Development**: Run with `streamlit run app.py` 2. **Streamlit Cloud**: Deploy directly from GitHub 3. **Docker Containers**: Package for deployment anywhere 4. **Cloud Providers**: Deploy to AWS, GCP, Azure, or other providers ### Installation Requirements ```bash pip install streamlit langchain python-dotenv ``` ## Troubleshooting & FAQ - **Which models are supported?**:br Any tools capable model available via APIpie's OpenAI-compatible endpoint. - **How do I persist my API key and endpoint?**:br Add the `export` lines to your shell profile or use a .env file. - **Can I use streaming with LangChain?**:br Yes, set `streaming=True` in your ChatOpenAI configuration and use the appropriate callbacks. For more, see the [LangChain.js GitHub](https://github.com/langchain-ai/langchainjs){rel=""nofollow""} or [LangChain Python GitHub](https://github.com/langchain-ai/langchain){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # LangGraph Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![LangGraph](https://apipie.ai/img/docs/integrations/LangGraph/wordmark_light.svg){width="600"} :: This guide will walk you through integrating LangGraph with APIpie, enabling you to build sophisticated, stateful, multi-actor applications that leverage a wide range of language models. ## What is [LangGraph](https://langchain-ai.github.io/langgraph/){rel=""nofollow""}? LangGraph is a powerful orchestration framework for building controllable agents and multi-agent workflows. It provides: - **Stateful Workflows**: Create complex, multi-step agent workflows with persistent state - **Multi-Agent Systems**: Design systems where multiple agents collaborate to solve tasks - **Human-in-the-Loop**: Incorporate human feedback and approval in your agent workflows - **Streaming Support**: First-class token-by-token streaming for real-time visibility - **Graph-Based Design**: Model agent behaviors as computational graphs for flexibility - **Low-Level Control**: Build custom agents without rigid abstractions - **Persistence**: Maintain state across execution for reliable long-running workflows By connecting LangGraph with APIpie, you gain access to a wide range of powerful language models for your agent applications while leveraging LangGraph's sophisticated orchestration capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with LangGraph. ### 4. Install LangGraph and Required Packages ```bash pip install langgraph langchain langchain_openai ``` You may need additional packages depending on your specific use case: ```bash # For retrieval-augmented agents pip install langchain-community # For human feedback pip install panel # For tracing and debugging pip install langsmith ``` ### 5. Configure LangGraph for APIpie LangGraph integrates with LangChain's LLM providers, which can be configured to use APIpie: ```python import os from langchain_openai import ChatOpenAI # Configure the LLM to use APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Create an instance of the LLM llm = ChatOpenAI( model="gpt-4o-mini", # Choose any model available on APIpie temperature=0.7 ) ``` --- ## Key Features - **Graph-Based Orchestration**: Build complex agent workflows as computational graphs - **State Management**: Maintain and update agent state across interactions - **Multi-Agent Collaboration**: Create teams of agents with different roles and capabilities - **Tool Integration**: Equip agents with tools to perform actions in the world - **Human Feedback**: Incorporate human input and oversight at critical decision points - **First-Class Streaming**: Stream agent reasoning and actions in real-time - **Checkpointing**: Persist agent state for reliable long-running tasks --- ## Example Workflows | Application Type | What LangGraph Helps You Build | | --------------------------- | ---------------------------------------------------------- | | Conversational Agents | Multi-turn conversational agents with memory and reasoning | | ReAct Agents | Agents that reason and act to solve complex tasks | | Research & Analysis | Multi-agent systems that collaborate on research problems | | Business Process Automation | Enterprise workflows with human approvals and oversight | | Supervised Agents | Agent systems with human-in-the-loop supervision | --- ## Using LangGraph with APIpie ### Basic ReAct Agent ```python import os from langgraph.prebuilt import create_react_agent from langchain_openai import ChatOpenAI # Configure APIpie credentials os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Create a tool that the agent can use def search(query: str) -> str: """Search for information on the internet.""" # Mock implementation - in a real app, you would call a search API if "weather" in query.lower(): return "It's currently 72°F and sunny in San Francisco." elif "population" in query.lower(): return "The population of the United States is approximately 332 million." else: return "No relevant information found." # Create a ReAct agent with APIpie as the LLM provider llm = ChatOpenAI(model="gpt-4o", temperature=0) agent = create_react_agent(llm, tools=[search]) # Use the agent to answer a question result = agent.invoke( {"messages": [{"role": "user", "content": "What's the weather in San Francisco?"}]} ) print(result["messages"][-1]["content"]) ``` ### Custom Multi-Step Agent Workflow ```python import os from typing import TypedDict, Annotated, List, Dict, Any from langchain_openai import ChatOpenAI from langchain_core.messages import HumanMessage, AIMessage from langchain_core.prompts import ChatPromptTemplate from langgraph.graph import StateGraph, END # Configure APIpie credentials os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Define the state for our workflow class AgentState(TypedDict): messages: List[Dict[str, Any]] summary: str # Set up the LLM with APIpie llm = ChatOpenAI(model="gpt-4o-mini", temperature=0.7) # Define the nodes in our graph # 1. Agent that processes user questions def process_query(state: AgentState) -> AgentState: """Process the user query and generate a response.""" # Get the last message last_message = state["messages"][-1] # Only process user messages if last_message["role"] != "user": return state # Create a prompt prompt = ChatPromptTemplate.from_messages([ ("system", "You are a helpful assistant. Answer user questions accurately."), ("human", "{input}") ]) # Get response from LLM response = llm.invoke(prompt.format(input=last_message["content"])) # Add the AI message to the state new_messages = state["messages"].copy() new_messages.append({"role": "assistant", "content": response.content}) return {"messages": new_messages, "summary": state["summary"]} # 2. Agent that creates a summary of the conversation def summarize_conversation(state: AgentState) -> AgentState: """Create a summary of the conversation so far.""" # Extract the conversation conversation = "\n".join([f"{m['role']}: {m['content']}" for m in state["messages"]]) # Create a prompt for summarization prompt = ChatPromptTemplate.from_messages([ ("system", "Summarize the following conversation in a concise paragraph."), ("human", "{conversation}") ]) # Get summary from LLM summary = llm.invoke(prompt.format(conversation=conversation)) return {"messages": state["messages"], "summary": summary.content} # 3. Routing function to decide when to summarize or end def should_summarize(state: AgentState) -> str: """Decide whether to summarize the conversation or continue.""" # Count message pairs (user + assistant) message_pairs = len([m for m in state["messages"] if m["role"] == "assistant"]) # Summarize after every 3 exchanges if message_pairs > 0 and message_pairs % 3 == 0: return "summarize" else: return "continue" # Create the workflow graph workflow = StateGraph(AgentState) # Add nodes workflow.add_node("process", process_query) workflow.add_node("summarize", summarize_conversation) # Set up the edges workflow.add_edge("process", should_summarize) workflow.add_conditional_edges( "should_summarize", { "summarize": "summarize", "continue": END } ) workflow.add_edge("summarize", END) # Set the entry point workflow.set_entry_point("process") # Compile the graph agent_workflow = workflow.compile() # Initialize the state initial_state = { "messages": [{"role": "user", "content": "Tell me about artificial intelligence."}], "summary": "" } # Run the workflow result = agent_workflow.invoke(initial_state) print("Final Messages:") for message in result["messages"]: print(f"{message['role']}: {message['content']}\n") print("Conversation Summary:") print(result["summary"]) ``` ### Multi-Agent Team with Human Approval ```python import os from typing import TypedDict, List, Dict, Any, Literal, Annotated from langchain_openai import ChatOpenAI from langchain_core.messages import HumanMessage, AIMessage from langchain_core.prompts import ChatPromptTemplate from langgraph.graph import StateGraph, END from langgraph.prebuilt import ToolExecutor, ToolInvocation # Configure APIpie credentials os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_API_BASE"] = "https://apipie.ai/v1" # Define our state class TeamState(TypedDict): messages: List[Dict[str, Any]] sender: str next: str task: str solution: str approval: bool # Define our tools def search_web(query: str) -> str: """Search the web for information.""" # Mock implementation return f"Results for '{query}': Found relevant information about {query}." def calculate(expression: str) -> str: """Calculate a mathematical expression.""" try: return str(eval(expression)) except: return "Error evaluating expression." # Set up tool executor tools = [search_web, calculate] tool_executor = ToolExecutor(tools) # Create our team of agents using APIpie model = ChatOpenAI(model="gpt-4o", temperature=0.7) # Define the system prompts for different roles researcher_prompt = ChatPromptTemplate.from_messages([ ("system", """You are a research specialist. Your job is to gather information to help solve problems. Use the search_web tool to find information related to the task. Be thorough and provide detailed results."""), ("human", "{task}") ]) analyst_prompt = ChatPromptTemplate.from_messages([ ("system", """You are an analyst. Your job is to analyze information and perform calculations when needed. Use the calculate tool for mathematical operations. Provide detailed analysis based on the information provided."""), ("human", "{task}\n\nResearch information: {research}") ]) solution_architect_prompt = ChatPromptTemplate.from_messages([ ("system", """You are a solution architect. Your job is to create a final solution based on research and analysis. Create a comprehensive and clear solution that addresses the original task. Be creative but practical."""), ("human", "{task}\n\nResearch information: {research}\n\nAnalysis: {analysis}") ]) # Define the agent functions def researcher(state: TeamState) -> TeamState: """Research agent that gathers information.""" if state["next"] != "researcher": return state # Get task information task = state["task"] # Research using the model research_response = model.invoke(researcher_prompt.format(task=task)) research_message = {"role": "assistant", "name": "researcher", "content": research_response.content} # Update state messages = state["messages"].copy() messages.append(research_message) return { "messages": messages, "sender": "researcher", "next": "analyst", "task": state["task"], "solution": state["solution"], "approval": state["approval"] } def analyst(state: TeamState) -> TeamState: """Analyst agent that analyzes information.""" if state["next"] != "analyst": return state # Get research information from the researcher research_messages = [m for m in state["messages"] if m.get("name") == "researcher"] research = research_messages[-1]["content"] if research_messages else "" # Analyze using the model analysis_response = model.invoke(analyst_prompt.format(task=state["task"], research=research)) analysis_message = {"role": "assistant", "name": "analyst", "content": analysis_response.content} # Update state messages = state["messages"].copy() messages.append(analysis_message) return { "messages": messages, "sender": "analyst", "next": "solution_architect", "task": state["task"], "solution": state["solution"], "approval": state["approval"] } def solution_architect(state: TeamState) -> TeamState: """Solution architect agent that creates the final solution.""" if state["next"] != "solution_architect": return state # Get research and analysis information research_messages = [m for m in state["messages"] if m.get("name") == "researcher"] analysis_messages = [m for m in state["messages"] if m.get("name") == "analyst"] research = research_messages[-1]["content"] if research_messages else "" analysis = analysis_messages[-1]["content"] if analysis_messages else "" # Create solution using the model solution_response = model.invoke(solution_architect_prompt.format( task=state["task"], research=research, analysis=analysis )) solution_message = {"role": "assistant", "name": "solution_architect", "content": solution_response.content} # Update state messages = state["messages"].copy() messages.append(solution_message) return { "messages": messages, "sender": "solution_architect", "next": "human_approval", "task": state["task"], "solution": solution_response.content, "approval": state["approval"] } def human_approval(state: TeamState) -> TeamState: """Simulate human approval (in a real app, this would wait for human input).""" if state["next"] != "human_approval": return state # Simulate approval (in a real application, this would be a UI for human input) # Here we're just automatically approving approval_message = {"role": "human", "name": "supervisor", "content": "Solution approved."} # Update state messages = state["messages"].copy() messages.append(approval_message) return { "messages": messages, "sender": "human", "next": "end", "task": state["task"], "solution": state["solution"], "approval": True } # Create the graph team_graph = StateGraph(TeamState) # Add nodes team_graph.add_node("researcher", researcher) team_graph.add_node("analyst", analyst) team_graph.add_node("solution_architect", solution_architect) team_graph.add_node("human_approval", human_approval) # Add edges team_graph.add_edge("researcher", "analyst") team_graph.add_edge("analyst", "solution_architect") team_graph.add_edge("solution_architect", "human_approval") team_graph.add_edge("human_approval", END) # Set entry point team_graph.set_entry_point("researcher") # Compile the graph team_workflow = team_graph.compile() # Initialize state initial_state = { "messages": [], "sender": "user", "next": "researcher", "task": "Research the impact of artificial intelligence on healthcare and propose three innovative applications.", "solution": "", "approval": False } # Run the workflow result = team_workflow.invoke(initial_state) # Display the results print("Task:", result["task"]) print("\nFinal Solution:", result["solution"]) print("\nApproval Status:", "Approved" if result["approval"] else "Pending") ``` --- ## Troubleshooting & FAQ - **How do I debug my LangGraph workflows?**:br LangGraph integrates with LangSmith for tracing and debugging. You can also add print statements at each node to see the state evolution. - **Can I use streaming with APIpie and LangGraph?**:br Yes, streaming is fully supported. Use the `streaming=True` parameter when creating your LLM instances. - **How do I handle environment variables securely?**:br Store your API keys in environment variables and never expose them in your code or repositories. - **Can I persist agent state between sessions?**:br Yes, LangGraph provides checkpointing functionality that allows you to persist and restore state across sessions. - **How do I integrate human feedback?**:br LangGraph supports human-in-the-loop workflows through dedicated human intervention nodes and tools like Panel for UI elements. - **What's the difference between LangGraph and other agent frameworks?**:br LangGraph provides low-level orchestration primitives with graph-based control flow, making it more flexible for complex agent systems than rigid frameworks. For more information, see the [LangGraph documentation](https://langchain-ai.github.io/langgraph/){rel=""nofollow""}, [GitHub repository](https://github.com/langchain-ai/langgraph){rel=""nofollow""}, or [LangChain Academy](https://academy.langchain.com/courses/intro-to-langgraph){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # LlamaIndex Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![LlamaIndex](https://apipie.ai/img/docs/integrations/Llamaindex/LlamaSquareBlack.svg){height="125"} :: This guide will walk you through integrating LlamaIndex with APIpie, enabling you to build powerful RAG (Retrieval Augmented Generation) applications that connect your custom data sources to various LLMs through a unified interface. ## What is [LlamaIndex](https://www.llamaindex.ai/){rel=""nofollow""}? LlamaIndex is a comprehensive data framework for connecting custom data to LLMs. It provides tools for: - **Data Ingestion**: Connect to various data sources through built-in connectors - **Data Indexing**: Structure your data for efficient retrieval - **Retrieval**: Extract relevant context for queries - **Synthesis**: Generate accurate responses augmented with retrieved context - **Evaluation**: Assess and improve RAG system performance - **Agent Orchestration**: Build complex LLM agents with data access By connecting LlamaIndex with APIpie, you gain access to a wide range of powerful language models while leveraging LlamaIndex's sophisticated data management capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with LlamaIndex. ### 4. Install LlamaIndex Install LlamaIndex core and required packages: ```bash pip install llama-index-core pip install llama-index-llms-openai # For OpenAI-compatible endpoints like APIpie ``` For advanced use cases, you may need additional packages: ```bash pip install llama-index-embeddings-openai # For embeddings pip install llama-index-vector-stores-qdrant # For Qdrant vector store # or other integrations as needed ``` ### 5. Configure LlamaIndex for APIpie Create a custom LLM configuration that points to APIpie: ```python import os from llama_index.llms.openai import OpenAI from llama_index.core import Settings # Configure APIpie as the LLM provider api_key = "your-apipie-api-key" apipie_llm = OpenAI( api_key=api_key, base_url="https://apipie.ai/v1", model="gpt-4o-mini", # You can use any model available on APIpie temperature=0.1, ) # Set as the default LLM Settings.llm = apipie_llm ``` --- ## Key Features - **Multiple Data Connector Options**: Connect to APIs, PDFs, CSVs, SQL databases, websites, and more - **Flexible Data Indexing**: Create vector stores, summaries, keyword indices, and knowledge graphs - **Advanced Retrieval**: Implement BM25, hybrid search, or re-ranking for improved results - **Query Planning**: Break down complex queries into sub-questions for comprehensive answers - **Caching**: Optimize performance and reduce API costs - **Evaluation Framework**: Assess and fine-tune your RAG systems --- ## Example Workflows | Application Type | What LlamaIndex Helps You Build | | ------------------------ | ---------------------------------------------------------- | | Document Q\&A | Systems that answer questions about specific documents | | Knowledge Bases | Comprehensive knowledge systems from multiple data sources | | Research Assistants | Tools that analyze and synthesize information | | Data Analysis | Systems that query and analyze structured data | | Multi-Agent Applications | Complex agent systems with coordinated data access | --- ## Using LlamaIndex with APIpie ### Basic Document Q\&A ```python import os from llama_index.llms.openai import OpenAI from llama_index.core import ( VectorStoreIndex, SimpleDirectoryReader, Settings, ) # Configure APIpie api_key = "your-apipie-api-key" apipie_llm = OpenAI( api_key=api_key, base_url="https://apipie.ai/v1", model="gpt-4o-mini", # You can use any model available on APIpie temperature=0.1, ) # Set as the default LLM Settings.llm = apipie_llm # Load your documents documents = SimpleDirectoryReader("./data").load_data() # Create an index from the documents index = VectorStoreIndex.from_documents(documents) # Create a query engine query_engine = index.as_query_engine() # Query your data response = query_engine.query("What is the main topic discussed in these documents?") print(response) ``` ### Advanced RAG with Custom Embeddings ```python import os from llama_index.llms.openai import OpenAI from llama_index.embeddings.openai import OpenAIEmbedding from llama_index.core import ( VectorStoreIndex, SimpleDirectoryReader, Settings, ServiceContext, ) # Configure APIpie for LLM api_key = "your-apipie-api-key" apipie_llm = OpenAI( api_key=api_key, base_url="https://apipie.ai/v1", model="gpt-4o", temperature=0.1, ) # Configure APIpie for embeddings apipie_embed_model = OpenAIEmbedding( api_key=api_key, base_url="https://apipie.ai/v1", model_name="text-embedding-3-large", embed_batch_size=100, ) # Set the default models Settings.llm = apipie_llm Settings.embed_model = apipie_embed_model # Load your documents documents = SimpleDirectoryReader("./data").load_data() # Create an index with the custom settings index = VectorStoreIndex.from_documents(documents) # Create a query engine with more advanced settings query_engine = index.as_query_engine( similarity_top_k=5, # Retrieve top 5 most similar chunks streaming=True, # Enable streaming responses ) # Query your data response = query_engine.query( "Provide a detailed summary of these documents and their key insights." ) print(response) ``` ### Using a Persistent Vector Store ```python import os from llama_index.llms.openai import OpenAI from llama_index.vector_stores.qdrant import QdrantVectorStore from llama_index.core import ( VectorStoreIndex, SimpleDirectoryReader, Settings, StorageContext, ) import qdrant_client # Configure APIpie api_key = "your-apipie-api-key" apipie_llm = OpenAI( api_key=api_key, base_url="https://apipie.ai/v1", model="gpt-4o-mini", ) # Set as the default LLM Settings.llm = apipie_llm # Create a Qdrant client (local or cloud) client = qdrant_client.QdrantClient( location=":memory:", # Use a real URL for production ) # Create a QdrantVectorStore vector_store = QdrantVectorStore( client=client, collection_name="documents", ) # Create a storage context storage_context = StorageContext.from_defaults(vector_store=vector_store) # Load your documents documents = SimpleDirectoryReader("./data").load_data() # Create an index with the custom vector store index = VectorStoreIndex.from_documents( documents, storage_context=storage_context, ) # Create a query engine query_engine = index.as_query_engine() # Query your data response = query_engine.query("What are the main points in these documents?") print(response) ``` ### Building an Agent with Data Access ```python import os from llama_index.llms.openai import OpenAI from llama_index.core import ( VectorStoreIndex, SimpleDirectoryReader, Settings, ) from llama_index.core.tools import QueryEngineTool, ToolMetadata from llama_index.core.query_engine import SubQuestionQueryEngine from llama_index.core.callbacks import CallbackManager, ConsoleCallbackHandler # Configure callbacks for logging Settings.callback_manager = CallbackManager([ConsoleCallbackHandler()]) # Configure APIpie api_key = "your-apipie-api-key" apipie_llm = OpenAI( api_key=api_key, base_url="https://apipie.ai/v1", model="gpt-4o", # Using a more capable model for the agent temperature=0.1, ) # Set as the default LLM Settings.llm = apipie_llm # Load different document sets financial_docs = SimpleDirectoryReader("./financial_data").load_data() product_docs = SimpleDirectoryReader("./product_data").load_data() # Create indices for each document set financial_index = VectorStoreIndex.from_documents(financial_docs) product_index = VectorStoreIndex.from_documents(product_docs) # Create query engines for each index financial_engine = financial_index.as_query_engine() product_engine = product_index.as_query_engine() # Create tools from the query engines tools = [ QueryEngineTool( query_engine=financial_engine, metadata=ToolMetadata( name="financial_data", description="Provides information about financial statements, revenue, and business performance", ), ), QueryEngineTool( query_engine=product_engine, metadata=ToolMetadata( name="product_data", description="Provides information about products, features, and specifications", ), ), ] # Create a sub-question query engine that can route to the appropriate tool query_engine = SubQuestionQueryEngine.from_defaults( query_engine_tools=tools, verbose=True, ) # Query across both datasets response = query_engine.query( "Compare the financial performance of our top-selling product to the overall company results last quarter." ) print(response) ``` --- ## Troubleshooting & FAQ - **Which models are supported?**:br Any OpenAI-compatible model available via APIpie's endpoint. - **How do I handle environment variables securely?**:br Store your API keys in environment variables or use a secure environment management tool. Never commit API keys to repositories. - **How can I optimize token usage?**:br LlamaIndex provides several mechanisms to reduce token usage, including chunking strategies, text splitting parameters, and caching. You can adjust the `chunk_size` and `chunk_overlap` parameters in the document loading process. - **What if I need to handle very large datasets?**:br For large datasets, consider using a persistent vector database like Pinecone, Weaviate, or Qdrant. LlamaIndex provides integrations with many vector stores. - **How do I debug retrieval issues?**:br Use the `verbose=True` parameter when creating query engines and enable the console callback handler to see detailed logs of the retrieval process. - **Can I use custom embedding models?**:br Yes, LlamaIndex supports many embedding models. You can configure custom embedding models through the appropriate integration packages and the `Settings.embed_model` configuration. For more information, see the [LlamaIndex documentation](https://docs.llamaindex.ai/){rel=""nofollow""} or the [GitHub repository](https://github.com/run-llama/llama_index){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # OpenAI Agents Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![OpenAI Agents](https://apipie.ai/img/docs/openai-white.png){height="125"} :: This guide will walk you through integrating the OpenAI Agents SDK with APIpie, enabling you to build powerful multi-agent workflows that leverage a wide range of language models. ## What is [OpenAI Agents SDK](https://github.com/openai/openai-agents-python){rel=""nofollow""}? OpenAI Agents SDK is a lightweight yet powerful framework for building multi-agent workflows. It provides a set of tools and components for: - **Agent Orchestration**: Create and manage multiple specialized agents - **Agent Handoffs**: Seamlessly transfer control between agents - **Guardrails**: Configure safety checks for input and output validation - **Tracing**: Built-in tracking and debugging of agent runs - **Tool Integration**: Enable agents to use functions and external APIs - **Provider Agnosticism**: Works with OpenAI and 100+ other LLMs By connecting OpenAI Agents SDK to APIpie, you gain access to a wide range of powerful language models while leveraging the SDK's sophisticated agent orchestration capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with OpenAI Agents. ### 4. Install OpenAI Agents SDK ```bash pip install openai-agents ``` For voice support (optional): ```bash pip install 'openai-agents[voice]' ``` ### 5. Configure OpenAI Agents SDK for APIpie Set up the environment variables for the APIpie integration: ```python import os # Set APIpie as the provider os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_BASE_URL"] = "https://apipie.ai/v1" # Optional: Configure default model os.environ["OPENAI_MODEL"] = "gpt-4o" # Use any model available on APIpie ``` --- ## Key Features - **Multi-Agent Orchestration**: Create specialized agents and coordinate their interactions - **Seamless Handoffs**: Transfer control between agents based on needs and expertise - **Configurable Guardrails**: Set safety checks for inputs and outputs - **Comprehensive Tracing**: Debug and optimize agent workflows - **Tool Integration**: Equip agents with functions to interact with external systems - **Model Flexibility**: Works with any LLM accessible via OpenAI-compatible API --- ## Example Workflows | Application Type | What OpenAI Agents Helps You Build | | ------------------------------ | -------------------------------------------------------- | | Multi-Specialist Systems | Workflows with specialized agents for different domains | | Human-in-the-Loop Applications | Systems with human review and intervention points | | Complex Reasoning Chains | Applications requiring multi-step reasoning and planning | | Enterprise Workflows | Business processes with multiple specialized steps | | Safety-Critical Applications | Systems with built-in guardrails and validation | --- ## Using OpenAI Agents SDK with APIpie ### Basic Single Agent Example ```python import os from agents import Agent, Runner # Configure APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_BASE_URL"] = "https://apipie.ai/v1" # Create a simple agent agent = Agent( name="Assistant", instructions="You are a helpful assistant specializing in Python programming.", model="gpt-4o", # Use any model available on APIpie ) # Run the agent result = Runner.run_sync( agent, "Explain how to use list comprehensions in Python with some examples." ) print(result.final_output) ``` ### Multi-Agent System with Handoffs ```python import os import asyncio from agents import Agent, Runner # Configure APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_BASE_URL"] = "https://apipie.ai/v1" # Create specialized agents python_agent = Agent( name="Python Expert", instructions="You are an expert Python programmer. Provide detailed, technically accurate information about Python programming.", model="gpt-4o", ) javascript_agent = Agent( name="JavaScript Expert", instructions="You are an expert JavaScript programmer. Provide detailed, technically accurate information about JavaScript programming.", model="gpt-4o", ) # Create a triage agent that can hand off to specialists triage_agent = Agent( name="Programming Triage", instructions="Determine if the user is asking about Python or JavaScript and hand off to the appropriate expert agent.", model="gpt-4o-mini", # Use a lighter model for triage handoffs=[python_agent, javascript_agent], ) async def main(): # Run the agent system result = await Runner.run( triage_agent, "What's the difference between list comprehensions in Python and array methods in JavaScript?" ) print(result.final_output) if __name__ == "__main__": asyncio.run(main()) ``` ### Using Function Tools ```python import os import asyncio from agents import Agent, Runner, function_tool # Configure APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_BASE_URL"] = "https://apipie.ai/v1" # Define a function tool for weather information @function_tool def get_weather(city: str, country: str = "US") -> str: """Get the current weather for a city. Args: city: The name of the city country: The country code (default: US) Returns: Current weather information """ # In a real implementation, you would call a weather API return f"The weather in {city}, {country} is currently sunny and 72°F." # Create an agent with the weather tool weather_agent = Agent( name="Weather Assistant", instructions="You help users get weather information for different locations.", model="gpt-4o-mini", tools=[get_weather], ) async def main(): result = await Runner.run( weather_agent, "What's the weather like in Tokyo, Japan?" ) print(result.final_output) if __name__ == "__main__": asyncio.run(main()) ``` ### Adding Guardrails ```python import os import asyncio from agents import Agent, Runner from agents.guardrails import InputGuardrail, OutputGuardrail # Configure APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_BASE_URL"] = "https://apipie.ai/v1" # Define guardrails class ProfanityInputGuardrail(InputGuardrail): async def validate(self, input_text: str) -> bool: profanity_list = ["bad_word1", "bad_word2"] # Define your list of prohibited words for word in profanity_list: if word in input_text.lower(): self.failure_reason = f"Input contains prohibited word: {word}" return False return True class FactualOutputGuardrail(OutputGuardrail): async def validate(self, output_text: str) -> bool: if "definitely" in output_text.lower() and "always" in output_text.lower(): self.failure_reason = "Output contains overgeneralizations" return False return True # Create an agent with guardrails agent = Agent( name="Guarded Assistant", instructions="You provide helpful information about science topics.", model="gpt-4o", input_guardrails=[ProfanityInputGuardrail()], output_guardrails=[FactualOutputGuardrail()], ) async def main(): try: result = await Runner.run( agent, "Tell me about the solar system." ) print(result.final_output) except Exception as e: print(f"Guardrail triggered: {e}") if __name__ == "__main__": asyncio.run(main()) ``` ### Using Structured Output Types ```python import os import asyncio from typing import List from pydantic import BaseModel from agents import Agent, Runner # Configure APIpie os.environ["OPENAI_API_KEY"] = "your-apipie-api-key" os.environ["OPENAI_BASE_URL"] = "https://apipie.ai/v1" # Define a structured output type class MovieRecommendation(BaseModel): title: str year: int director: str genre: str description: str class MovieRecommendations(BaseModel): recommendations: List[MovieRecommendation] reasoning: str # Create an agent with structured output movie_agent = Agent( name="Movie Recommender", instructions="You recommend movies based on user preferences.", model="gpt-4o", output_type=MovieRecommendations, ) async def main(): result = await Runner.run( movie_agent, "Recommend three sci-fi movies similar to Interstellar." ) # Access structured data for i, movie in enumerate(result.final_output.recommendations, 1): print(f"Recommendation {i}:") print(f" Title: {movie.title}") print(f" Year: {movie.year}") print(f" Director: {movie.director}") print(f" Genre: {movie.genre}") print(f" Description: {movie.description}") print(f"\nReasoning: {result.final_output.reasoning}") if __name__ == "__main__": asyncio.run(main()) ``` --- ## Troubleshooting & FAQ - **How do I configure specific models for different agents?**:br Set the `model` parameter when creating each agent: `Agent(name="Agent", model="gpt-4o")` - **Can I use custom prompts with OpenAI Agents?**:br Yes, use the `instructions` parameter to define your agent's behavior. For more complex prompting, use multiple agents with specialized instructions. - **How do I debug agent behavior?**:br OpenAI Agents SDK includes built-in tracing. You can also use external tracing processors like Logfire, AgentOps, Braintrust, Scorecard, or Keywords AI. - **What if my handoff logic doesn't work as expected?**:br Review your triage agent's instructions and ensure they clearly explain when to hand off to which agent. You can also implement explicit logic by using function tools. - **Are there limits to how many agents I can chain together?**:br While there's no strict limit, each handoff adds latency. The `max_turns` parameter can prevent infinite loops. - **How do I handle authentication for function tools?**:br Store API keys securely and use environment variables. Your function tool should handle authentication internally. For more information, see the [OpenAI Agents SDK documentation](https://openai.github.io/openai-agents-python/){rel=""nofollow""} or the [GitHub repository](https://github.com/openai/openai-agents-python){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Pixeltable Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"}s ![Pixeltable](https://apipie.ai/img/docs/integrations/Pixeltable/pixeltable-logo-large.png){height="125"} :: This guide will walk you through integrating Pixeltable with APIpie, enabling you to build powerful multimodal AI applications with declarative data infrastructure and access to a wide range of language models. ## What is [Pixeltable](https://pixeltable.com/){rel=""nofollow""}? Pixeltable is a declarative data infrastructure framework for multimodal AI apps that provides: - **Data Ingestion & Storage**: Work with images, videos, audio, documents, and structured data - **Transformation & Processing**: Apply Python functions or built-in operations automatically - **AI Model Integration**: Run inference (embeddings, object detection, LLMs) as part of data pipelines - **Indexing & Retrieval**: Create and manage vector indexes for semantic search - **Incremental Computation**: Only recompute what's necessary when data or code changes - **Versioning & Lineage**: Track data and schema changes for reproducibility By connecting Pixeltable with APIpie, you gain access to a wide range of powerful language models for your multimodal applications while leveraging Pixeltable's sophisticated data management capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with Pixeltable. ### 4. Install Pixeltable Install Pixeltable using pip: ```bash pip install pixeltable ``` For additional functionalities, you might need extra packages: ```bash # For working with embedding models pip install sentence-transformers # For object detection pip install pixeltable-yolox # For document processing pip install spacy python -m spacy download en_core_web_sm ``` ### 5. Configure Pixeltable for APIpie Pixeltable integrates with APIpie through built-in functions for OpenAI-compatible endpoints: ```python import pixeltable as pxt from pixeltable.functions import openai # Configure the OpenAI function with APIpie credentials # This configuration can be reused across your tables openai.configure( api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) ``` --- ## Key Features - **Unified Multimodal Interface**: Work with images, videos, audio, documents, and structured data through a consistent interface - **Declarative Computed Columns**: Define transformations that run automatically on new or updated data - **Built-in Vector Search**: Add embedding indexes for similarity search directly on tables/views - **On-the-Fly Data Views**: Create virtual tables using iterators for efficient processing - **Seamless AI Integration**: Built-in functions for various AI providers including APIpie - **Custom Python Functions**: Extend with User-Defined Functions (UDFs) - **Incremental Computation**: Save time and costs by only recomputing what's necessary --- ## Example Workflows | Application Type | What Pixeltable Helps You Build | | ------------------------ | -------------------------------------------------------- | | Multimodal RAG Systems | Apps that search and generate content from diverse media | | Computer Vision Apps | Applications performing detection and classification | | Content Generation | Systems that create and modify text and images | | Data Curation & Labeling | Tools for organizing and annotating multimodal datasets | | AI Evaluation Frameworks | Systems to test and benchmark model performance | --- ## Using Pixeltable with APIpie ### Basic Text Generation with APIpie ```python import pixeltable as pxt from pixeltable.functions import openai # Configure APIpie credentials openai.configure( api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) # Create a table for prompts and responses qa = pxt.create_table( 'qa_system', {'prompt': pxt.String}, if_exists='replace' ) # Add a computed column for LLM responses qa.add_computed_column( response=openai.chat_completions( model='gpt-4o-mini', # You can use any model available on APIpie messages=[{ 'role': 'user', 'content': qa.prompt }] ).choices[0].message.content ) # Insert a prompt and get a response qa.insert([ {'prompt': 'Explain quantum computing in simple terms.'} ]) # View the results print(qa.select(qa.prompt, qa.response).collect()) ``` ### Multimodal RAG with Pixeltable and APIpie ```python import pixeltable as pxt from pixeltable.functions import openai, huggingface from pixeltable.iterators import DocumentSplitter # Configure APIpie credentials openai.configure( api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) # Create a directory for organization pxt.create_dir("rag_system", if_exists="replace") # Create a document table and add some PDFs docs = pxt.create_table( 'rag_system.documents', {'doc': pxt.Document}, if_exists='replace' ) docs.insert([ {'doc': 'path/to/your/document.pdf'} ]) # Create chunks view with sentence-based splitting chunks = pxt.create_view( 'rag_system.chunks', docs, iterator=DocumentSplitter.create( document=docs.doc, separators='sentence' ) ) # Add embedding index for similarity search chunks.add_embedding_index( 'text', string_embed=huggingface.sentence_transformer.using( model_id='all-MiniLM-L6-v2' ) ) # Define query function for retrieval @pxt.query def get_relevant_context(query_text: str, limit: int = 3): sim = chunks.text.similarity(query_text) return chunks.order_by(sim, asc=False).limit(limit).select(chunks.text) # Create QA table that integrates context retrieval with APIpie qa = pxt.create_table( 'rag_system.qa', {'question': pxt.String}, if_exists='replace' ) # Add context retrieval as a computed column qa.add_computed_column( context=get_relevant_context(qa.question) ) # Construct prompt with retrieved context qa.add_computed_column( formatted_prompt=f""" Based on the following information, please answer the question. CONTEXT: {qa.context} QUESTION: {qa.question} ANSWER: """ ) # Generate response with APIpie qa.add_computed_column( answer=openai.chat_completions( model='gpt-4o', messages=[{ 'role': 'user', 'content': qa.formatted_prompt }] ).choices[0].message.content ) # Ask a question qa.insert([ {'question': 'What are the key concepts discussed in the document?'} ]) # View the result print(qa.select(qa.question, qa.answer).collect()) ``` ### Image Classification with Pixeltable and APIpie Vision ```python import pixeltable as pxt from pixeltable.functions import openai # Configure APIpie credentials openai.configure( api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) # Create an image table images = pxt.create_table( 'image_analysis', {'image': pxt.Image}, if_exists='replace' ) # Insert some sample images images.insert([ {'image': 'https://upload.wikimedia.org/wikipedia/commons/thumb/6/68/Orange_tabby_cat_sitting_on_fallen_leaves-Hisashi-01A.jpg/1920px-Orange_tabby_cat_sitting_on_fallen_leaves-Hisashi-01A.jpg'}, {'image': 'https://upload.wikimedia.org/wikipedia/commons/d/d5/Retriever_in_water.jpg'} ]) # Add a computed column for vision analysis using APIpie's vision model images.add_computed_column( description=openai.chat_completions( model='gpt-4-vision', messages=[ { 'role': 'user', 'content': [ {'type': 'text', 'text': 'Describe this image in detail.'}, { 'type': 'image_url', 'image_url': {'url': images.image} } ] } ] ).choices[0].message.content ) # View the results print(images.select(images.image, images.description).collect()) ``` ### Agentic Workflows with Tool Calling ```python import pixeltable as pxt from pixeltable.functions import openai import json # Configure APIpie credentials openai.configure( api_key="your-apipie-api-key", base_url="https://apipie.ai/v1" ) # Define a UDF for weather information (mock function) @pxt.udf def get_weather(location: str) -> str: """Get current weather for a location.""" # In a real app, you would call a weather API return f"The weather in {location} is sunny and 72°F." # Define a UDF for search (mock function) @pxt.udf def search_web(query: str) -> str: """Search the web for information.""" # In a real app, you would call a search API return f"Search results for: {query}..." # Register the UDFs as tools tools = pxt.tools(get_weather, search_web) # Create agent table agent = pxt.create_table( 'agent_system', {'user_request': pxt.String}, if_exists='replace' ) # Add a column for tool selection by the LLM agent.add_computed_column( tool_choice=openai.chat_completions( model='gpt-4o', messages=[ {'role': 'user', 'content': agent.user_request} ], tools=[t.to_openai_tool() for t in tools], tool_choice='auto' ) ) # Add a computed column to extract tool call information @pxt.udf def extract_tool_call(response) -> dict: """Extract tool call information from the LLM response.""" tool_calls = response.choices[0].message.tool_calls if not tool_calls: return {"tool": None, "args": None} call = tool_calls[0] return { "tool": call.function.name, "args": json.loads(call.function.arguments) } agent.add_computed_column( tool_info=extract_tool_call(agent.tool_choice) ) # Add a column to invoke the selected tool @pxt.udf def invoke_tool(tool_info: dict) -> str: """Invoke the selected tool with the provided arguments.""" if not tool_info or not tool_info["tool"]: return "No tool was selected." tool_name = tool_info["tool"] args = tool_info["args"] if tool_name == "get_weather": return get_weather(args["location"]) elif tool_name == "search_web": return search_web(args["query"]) return f"Unknown tool: {tool_name}" agent.add_computed_column( tool_result=invoke_tool(agent.tool_info) ) # Add a column for the final response agent.add_computed_column( final_response=openai.chat_completions( model='gpt-4o-mini', messages=[ {'role': 'user', 'content': agent.user_request}, {'role': 'assistant', 'content': f"I found this information: {agent.tool_result}"} ] ).choices[0].message.content ) # Test the agent agent.insert([ {'user_request': 'What\'s the weather like in Paris?'}, {'user_request': 'Tell me about quantum computing.'} ]) # View the results print(agent.select(agent.user_request, agent.tool_info, agent.final_response).collect()) ``` --- ## Troubleshooting & FAQ - **How do I update computed columns when data changes?**:br Pixeltable automatically recomputes values when data or UDFs change. To force a recomputation, use `table.invalidate_columns(['column_name'])`. - **How do I handle environment variables securely?**:br Store your API keys in environment variables and load them in your code. Pixeltable also supports configuration files for security. - **Can I use APIpie's routing capabilities with Pixeltable?**:br Yes, APIpie's routing features work seamlessly with Pixeltable. Simply specify the desired model when calling functions. - **How do I monitor costs and usage?**:br Pixeltable's incremental computation helps minimize API calls, but you should still monitor your APIpie dashboard for usage metrics. - **Can I cache API calls to reduce costs?**:br Yes, Pixeltable's computed columns implicitly cache results, only recomputing when inputs change. - **How can I scale to large datasets?**:br Pixeltable is designed to handle large datasets efficiently. For very large datasets, consider using Views with iterators for on-demand processing. For more information, see the [Pixeltable documentation](https://docs.pixeltable.com/){rel=""nofollow""} or the [GitHub repository](https://github.com/pixeltable/pixeltable){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} or [Pixeltable Discord](https://discord.gg/QPyqFYx2UN){rel=""nofollow""} for assistance. # PydanticAI Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![PydanticAI](https://apipie.ai/img/docs/integrations/PydanticAI/PydanticAI.svg){height="125"} :: This guide will walk you through integrating PydanticAI with APIpie, enabling you to build type-safe AI applications with structured outputs. ## What is PydanticAI? PydanticAI is a Python agent framework designed to make it easier to build production grade applications with Generative AI. Created by the team behind Pydantic, it brings the same level of type safety and validation to AI applications that Pydantic brings to Python data validation. Key features include: - Type-safe interfaces for working with LLMs - Structured validation of model outputs - Support for tools and agents - Dependency injection for modular design - Built-in support for popular LLM providers By connecting PydanticAI to APIpie, you unlock access to powerful language models while maintaining type safety and structured outputs. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with PydanticAI. ### 4. Install PydanticAI ```bash pip install 'pydantic-ai-slim[openai]' ``` ### 5. Configure PydanticAI for APIpie You can use PydanticAI with APIpie through its OpenAI-compatible interface: ```python from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIModel model = OpenAIModel( "openai/gpt-4o-mini", # or any other APIpie model base_url="https://apipie.ai/v1", api_key="your-APIpie-key-here", ) agent = Agent(model) result = await agent.run("Why is the sky blue?") print(result) ``` --- ## Key Features - **Type Safety**: Build AI applications that are statically type-checked - **Structured Outputs**: Guarantee that LLM responses follow your defined schemas - **Tool Framework**: Create tools that LLMs can use, with automatic parameter validation - **Dependency Injection**: Easily provide context and services to your agents - **Python-First Design**: Use familiar Python patterns and practices with AI --- ## Example Workflows | Application Type | What PydanticAI Helps You Build | | -------------------------- | -------------------------------------------------------------- | | Type-Safe Assistants | Assistants that return validated, structured data | | Function-Calling Agents | Agents that can invoke functions and validate their parameters | | Business Logic Integration | AI systems that safely integrate with your existing systems | | Multi-Agent Systems | Complex workflows with multiple specialized agents | | Data Processing Pipelines | Pipelines that extract, transform, and validate data with LLMs | --- ## Using PydanticAI with APIpie ### Basic Agent ```python from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIModel # Configure the model model = OpenAIModel( "openai/gpt-4o-mini", base_url="https://apipie.ai/v1", api_key="your-APIpie-key-here", ) # Create a simple agent agent = Agent( model, system_prompt="You are a helpful assistant. Be concise and clear." ) # Run the agent synchronously result = agent.run_sync("What are three interesting facts about the moon?") print(result.output) ``` ### Using Structured Output ```python from pydantic import BaseModel, Field from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIModel # Define a structured output model class MoonFacts(BaseModel): fact1: str = Field(description="First interesting fact about the moon") fact2: str = Field(description="Second interesting fact about the moon") fact3: str = Field(description="Third interesting fact about the moon") # Configure the model model = OpenAIModel( "openai/gpt-4o-mini", base_url="https://apipie.ai/v1", api_key="your-APIpie-key-here", ) # Create an agent with structured output facts_agent = Agent( model, output_type=MoonFacts, system_prompt="You are a space expert." ) # Run the agent result = facts_agent.run_sync("Tell me about the moon") # Access the structured data print(f"Fact 1: {result.output.fact1}") print(f"Fact 2: {result.output.fact2}") print(f"Fact 3: {result.output.fact3}") ``` --- ## Troubleshooting & FAQ - **Which models are supported?**:br Any model available via APIpie's OpenAI-compatible endpoint. - **How do I persist my API key and endpoint?**:br Use environment variables or a .env file to store your credentials securely. - **Can I use streaming with PydanticAI?**:br Yes, through the `agent.stream()` method. For more information, see the [PydanticAI documentation](https://ai.pydantic.dev){rel=""nofollow""} or the [GitHub repository](https://github.com/pydantic/pydantic-ai){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Microsoft Semantic Kernel Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"}s ![Microsoft Semantic Kernel](https://apipie.ai/img/docs/integrations/Microsoft/Microsoft_logo.svg){height="125"} :: This guide will walk you through integrating Microsoft Semantic Kernel with APIpie, enabling you to build intelligent AI agents and multi-agent systems with enterprise-ready capabilities. ## What is [Microsoft Semantic Kernel](https://learn.microsoft.com/en-us/semantic-kernel/overview/){rel=""nofollow""}? Microsoft Semantic Kernel is a model-agnostic SDK that empowers developers to build, orchestrate, and deploy AI agents and multi-agent systems. It provides: - **Agent Framework**: Build modular AI agents with access to tools/plugins, memory, and planning - **Multi-Agent Systems**: Orchestrate complex workflows with collaborating specialist agents - **Plugin Ecosystem**: Extend with native code functions, prompt templates, OpenAPI specs, or Model Context Protocol (MCP) - **Vector DB Support**: Seamlessly integrate with various vector databases for knowledge retrieval - **Multimodal Support**: Process text, vision, and audio inputs - **Process Framework**: Model complex business processes with a structured workflow approach - **Enterprise Ready**: Built for observability, security, and stable APIs By connecting Semantic Kernel to APIpie, you gain access to a wide range of powerful language models while leveraging Semantic Kernel's sophisticated agent orchestration capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with Semantic Kernel. ### 4. Install Semantic Kernel #### Python ```bash pip install semantic-kernel ``` #### .NET ```bash dotnet add package Microsoft.SemanticKernel dotnet add package Microsoft.SemanticKernel.Agents.Core ``` ### 5. Configure Semantic Kernel for APIpie Semantic Kernel supports OpenAI-compatible APIs, which makes it easy to integrate with APIpie. #### Python ```python import os import semantic_kernel as sk from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion # Set up the kernel kernel = sk.Kernel() # Add APIpie as a chat service api_key = os.environ.get("APIPIE_API_KEY", "your-apipie-key") service_id = "apipie-chat" kernel.add_service( OpenAIChatCompletion( service_id=service_id, ai_model_id="gpt-4o", # Use any model available on APIpie api_key=api_key, endpoint="https://apipie.ai/v1" ) ) ``` #### .NET ```csharp using Microsoft.SemanticKernel; using Microsoft.SemanticKernel.ChatCompletion; // Create a kernel builder var builder = Kernel.CreateBuilder(); // Add APIpie as the chat completion service builder.AddOpenAIChatCompletion( modelId: "gpt-4o", // Use any model available on APIpie apiKey: Environment.GetEnvironmentVariable("APIPIE_API_KEY") ?? "your-apipie-key", serviceId: "apipie-chat", endpoint: new Uri("https://apipie.ai/v1") ); // Build the kernel var kernel = builder.Build(); ``` --- ## Key Features - **Model Flexibility**: Connect to any LLM with built-in support for OpenAI-compatible APIs - **Agent Building Blocks**: Create agents with specific instructions, tools, and memory - **Plugin System**: Extend functionality with native code functions, prompt templates, or APIs - **Memory & Embeddings**: Store and retrieve contextual information for more effective agents - **Planning Capabilities**: Enable agents to break down complex tasks into steps - **Streaming Support**: Process responses as they're generated for responsive applications - **Multi-Modal Support**: Handle text, images, and other data types in agent interactions - **Cross-Platform**: Available for Python, .NET, and Java developers --- ## Example Workflows | Application Type | What Semantic Kernel Helps You Build | | ------------------------- | --------------------------------------------------------------------- | | Conversational Assistants | Chatbots with memory, specialized knowledge, and tool access | | Customer Support Systems | Agents that handle different categories of support inquiries | | Research & Analysis Tools | Multi-agent systems that collect, analyze, and synthesize data | | Enterprise Workflows | Complex business processes with human-in-the-loop capabilities | | Document Processing | Systems that understand, extract, and generate content from documents | --- ## Using Semantic Kernel with APIpie ### Basic Chat Agent (Python) ```python import os import asyncio from semantic_kernel.agents import ChatCompletionAgent from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion async def main(): # Configure the chat service with APIpie chat_service = OpenAIChatCompletion( ai_model_id="gpt-4o", # Use any model available on APIpie api_key=os.environ.get("APIPIE_API_KEY", "your-apipie-key"), endpoint="https://apipie.ai/v1" ) # Create a simple chat agent agent = ChatCompletionAgent( service=chat_service, name="APIpie-Assistant", instructions="You are a helpful assistant that provides concise and accurate information.", ) # Get a response to a user message response = await agent.get_response(messages="What is the relationship between AI and machine learning?") print(response.content) if __name__ == "__main__": asyncio.run(main()) ``` ### Basic Chat Agent (.NET) ```csharp using System; using System.Threading.Tasks; using Microsoft.SemanticKernel; using Microsoft.SemanticKernel.Agents; class Program { static async Task Main() { // Create a kernel builder var builder = Kernel.CreateBuilder(); // Add APIpie as the chat completion service builder.AddOpenAIChatCompletion( modelId: "gpt-4o", // Use any model available on APIpie apiKey: Environment.GetEnvironmentVariable("APIPIE_API_KEY") ?? "your-apipie-key", endpoint: new Uri("https://apipie.ai/v1") ); // Build the kernel var kernel = builder.Build(); // Create a chat agent var agent = new ChatCompletionAgent() { Name = "APIpie-Assistant", Instructions = "You are a helpful assistant that provides concise and accurate information.", Kernel = kernel, }; // Get a response to a user message await foreach (var response in agent.InvokeAsync("What is the relationship between AI and machine learning?")) { Console.WriteLine(response.Message?.Content); } } } ``` ### Agent with Custom Tools (Python) ```python import os import asyncio from typing import Annotated from semantic_kernel.agents import ChatCompletionAgent from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion from semantic_kernel.functions import kernel_function, KernelArguments # Define a custom plugin for weather information class WeatherPlugin: @kernel_function(description="Get the current weather for a given location.") def get_weather(self, location: Annotated[str, "The city and country"]) -> Annotated[str, "The current weather information"]: # In a real implementation, this would call a weather API return f"The weather in {location} is currently sunny with a temperature of 25°C (77°F)." @kernel_function(description="Get the weather forecast for the next few days.") def get_forecast(self, location: Annotated[str, "The city and country"], days: Annotated[int, "Number of days to forecast"] = 3) -> Annotated[str, "The weather forecast"]: # In a real implementation, this would call a weather API return f"The {days}-day forecast for {location} shows sunny conditions with temperatures ranging from 22°C to 28°C." async def main(): # Configure the chat service with APIpie chat_service = OpenAIChatCompletion( ai_model_id="gpt-4o", # Use any model available on APIpie api_key=os.environ.get("APIPIE_API_KEY", "your-apipie-key"), endpoint="https://apipie.ai/v1" ) # Create a chat agent with the weather plugin agent = ChatCompletionAgent( service=chat_service, name="Weather-Assistant", instructions="You are a helpful assistant that provides weather information when asked.", plugins=[WeatherPlugin()], ) # Get a response to a user message response = await agent.get_response(messages="What's the weather like in Tokyo today?") print(response.content) if __name__ == "__main__": asyncio.run(main()) ``` ### Agent with Custom Tools (.NET) ```csharp using System; using System.ComponentModel; using System.Threading.Tasks; using Microsoft.SemanticKernel; using Microsoft.SemanticKernel.Agents; class Program { static async Task Main() { // Create a kernel builder var builder = Kernel.CreateBuilder(); // Add APIpie as the chat completion service builder.AddOpenAIChatCompletion( modelId: "gpt-4o", // Use any model available on APIpie apiKey: Environment.GetEnvironmentVariable("APIPIE_API_KEY") ?? "your-apipie-key", endpoint: new Uri("https://apipie.ai/v1") ); // Build the kernel var kernel = builder.Build(); // Add the weather plugin kernel.Plugins.Add(KernelPluginFactory.CreateFromType("WeatherPlugin")); // Create a chat agent with the weather plugin var agent = new ChatCompletionAgent() { Name = "Weather-Assistant", Instructions = "You are a helpful assistant that provides weather information when asked.", Kernel = kernel, Arguments = new KernelArguments({ FunctionChoiceBehavior = FunctionChoiceBehavior.Auto() }) }; // Get a response to a user message await foreach (var response in agent.InvokeAsync("What's the weather like in Tokyo today?")) { Console.WriteLine(response.Message?.Content); } } } // Define a custom plugin for weather information public class WeatherPlugin { [KernelFunction, Description("Get the current weather for a given location.")] public string GetWeather([Description("The city and country")] string location) { // In a real implementation, this would call a weather API return $"The weather in {location} is currently sunny with a temperature of 25°C (77°F)."; } [KernelFunction, Description("Get the weather forecast for the next few days.")] public string GetForecast( [Description("The city and country")] string location, [Description("Number of days to forecast")] int days = 3) { // In a real implementation, this would call a weather API return $"The {days}-day forecast for {location} shows sunny conditions with temperatures ranging from 22°C to 28°C."; } } ``` ### Multi-Agent System with Handoffs (Python) ```python import os import asyncio from semantic_kernel.agents import ChatCompletionAgent from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion async def main(): # Configure the chat service with APIpie chat_service = OpenAIChatCompletion( ai_model_id="gpt-4o", # Use any model available on APIpie api_key=os.environ.get("APIPIE_API_KEY", "your-apipie-key"), endpoint="https://apipie.ai/v1" ) # Create specialized agents billing_agent = ChatCompletionAgent( service=chat_service, name="BillingAgent", instructions="You handle billing issues like charges, payment methods, cycles, fees, discrepancies, and payment failures.", ) refund_agent = ChatCompletionAgent( service=chat_service, name="RefundAgent", instructions="You assist users with refund inquiries, including eligibility, policies, processing, and status updates.", ) # Create a triage agent that can delegate to specialized agents triage_agent = ChatCompletionAgent( service=chat_service, name="TriageAgent", instructions="Evaluate user requests and forward them to BillingAgent or RefundAgent for targeted assistance. Provide the full answer to the user containing any information from the agents.", plugins=[billing_agent, refund_agent], ) # Process a user request user_query = "I was charged twice for my subscription last month and need a refund." response = await triage_agent.get_response(messages=user_query) print(f"User: {user_query}\n\nAgent: {response.content}") if __name__ == "__main__": asyncio.run(main()) ``` ### Using Memory for Context-Aware Responses (Python) ```python import os import asyncio from semantic_kernel import Kernel from semantic_kernel.agents import ChatCompletionAgent from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion, OpenAITextEmbedding from semantic_kernel.memory import VolatileMemoryStore async def main(): # Create a kernel kernel = Kernel() # Configure chat service with APIpie chat_service = OpenAIChatCompletion( ai_model_id="gpt-4o", # Use any model available on APIpie api_key=os.environ.get("APIPIE_API_KEY", "your-apipie-key"), endpoint="https://apipie.ai/v1" ) # Configure embedding service with APIpie embedding_service = OpenAITextEmbedding( ai_model_id="text-embedding-3-large", # Use any embedding model available on APIpie api_key=os.environ.get("APIPIE_API_KEY", "your-apipie-key"), endpoint="https://apipie.ai/v1" ) # Add services to the kernel kernel.add_service(chat_service) kernel.add_service(embedding_service) # Set up memory memory_store = VolatileMemoryStore() kernel.register_memory(memory_store) # Add some memories await kernel.memory.save_information_async( collection="user_data", id="user_preferences", text="The user prefers vegetarian food and enjoys hiking on weekends." ) await kernel.memory.save_information_async( collection="user_data", id="user_location", text="The user lives in Seattle, Washington." ) # Create a chat agent with access to memory agent = ChatCompletionAgent( service=chat_service, name="ContextAwareAssistant", instructions="""You are a helpful assistant that provides personalized responses based on what you know about the user. Before answering, check if you have relevant information about the user that could help personalize your response.""", kernel=kernel, ) # Query the memory and include relevant context in the response user_query = "Can you recommend some activities for this weekend?" # Retrieve relevant memories memories = await kernel.memory.search_async( collection="user_data", query=user_query, limit=5 ) # Build context from memories context = "User information:\n" for memory in memories: context += f"- {memory.text}\n" # Create a message with the context and query message = f"{context}\n\nUser query: {user_query}" # Get a response from the agent response = await agent.get_response(messages=message) print(f"User: {user_query}\n\nAgent: {response.content}") if __name__ == "__main__": asyncio.run(main()) ``` --- ## Troubleshooting & FAQ - **How do I handle authentication with APIpie?**:br Store your API key securely in environment variables and pass it to the Semantic Kernel configuration. Never hardcode API keys in your code. - **Which models work best with Semantic Kernel?**:br While most models available on APIpie will work with Semantic Kernel, models with strong tool use/function calling capabilities (like GPT-4o or Claude 3 Opus) work best for agent-based applications. - **How do I debug agent behavior?**:br Enable logging in Semantic Kernel to see detailed information about each step the agent takes. For Python, use the standard logging module; for .NET, use ILogger. - **Can I mix and match different LLM providers?**:br Yes, Semantic Kernel allows you to configure different services for different purposes. For example, you might use APIpie for chat completions but another provider for embeddings. - **How do I handle rate limits?**:br Implement retry logic in your application to handle rate limits. Semantic Kernel does not have built-in retry mechanisms, so you'll need to handle this in your code. - **Does Semantic Kernel support streaming?**:br Yes, Semantic Kernel supports streaming responses. Use the streaming methods provided by the library to process responses as they're generated. For more information, see the [Semantic Kernel documentation](https://learn.microsoft.com/en-us/semantic-kernel/overview/){rel=""nofollow""} or the [GitHub repository](https://github.com/microsoft/semantic-kernel){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Smolagents Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"}s ![Smolagents](https://apipie.ai/img/docs/integrations/Smolagents/smolagents.png){height="125"} :: This guide will walk you through integrating Smolagents with APIpie, enabling you to build powerful AI agents that leverage code execution for more complex reasoning and tool use. ## What is [Smolagents](https://huggingface.co/docs/smolagents/index){rel=""nofollow""}? Smolagents is a minimalist, yet powerful library for building AI agents that think in code. It offers: - **Code-First Agents**: Agents that express their reasoning and actions as Python code (rather than JSON or natural language) - **Minimalist Design**: Core agent logic in \~1,000 lines of code, with minimal abstractions - **Sandboxed Execution**: Secure code execution with E2B or Docker sandboxing - **Multi-Modal Support**: Handle text, images, audio, and video inputs - **Flexible Tool Integration**: Use tools from various ecosystems or create your own - **Model Agnosticism**: Works with any LLM provider By integrating Smolagents with APIpie, you can leverage APIpie's diverse model selection while building powerful, code-native agents that can reason through complex problems step by step. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with Smolagents. ### 4. Install Smolagents ```bash pip install smolagents ``` For sandboxed execution (optional but recommended): ```bash # For E2B sandboxing pip install e2b # Or for Docker-based sandboxing # Make sure Docker is installed on your system ``` ### 5. Configure Smolagents with APIpie ```python import os from smolagents import CodeAgent, LiteLLMModel # Configure the LLM using APIpie model = LiteLLMModel( model_id="apipie/gpt-4o", # Use any APIpie model api_key=os.environ.get("APIPIE_API_KEY"), api_base="https://apipie.ai/v1" ) # Create a simple agent agent = CodeAgent(model=model) ``` --- ## Key Features - **Code as Actions**: Agents express reasoning and actions as Python code, enabling more complex logic - **Tool Integration**: Easy creation and use of custom tools with typed inputs and outputs - **Secure Execution**: Multiple options for secure code execution (local, E2B, Docker) - **Model Flexibility**: Works with any LLM accessible through InferenceClient, OpenAI API, or LiteLLM - **Modularity**: Build multi-agent systems with specialized agents working together - **Hub Integration**: Share and reuse tools via Hugging Face Hub --- ## Example Workflows | Application Type | What Smolagents Helps You Build | | ----------------------- | -------------------------------------------------------------- | | Research Assistants | Agents that can search, analyze, and synthesize information | | Data Analysis | Code agents that can process, visualize, and interpret data | | Web Browsing Agents | Agents that can navigate and extract information from the web | | Multi-Agent Systems | Orchestrate multiple specialized agents to solve complex tasks | | Tool-Using Applications | Apps that dynamically select and use appropriate tools | --- ## Using Smolagents with APIpie ### Basic Agent with Web Search ```python import os from smolagents import CodeAgent, DuckDuckGoSearchTool, LiteLLMModel # Configure with APIpie model = LiteLLMModel( model_id="apipie/gpt-4o-mini", api_key=os.environ.get("APIPIE_API_KEY"), api_base="https://apipie.ai/v1" ) # Create an agent with search capability agent = CodeAgent( tools=[DuckDuckGoSearchTool()], model=model ) # Run the agent result = agent.run("How many seconds would it take for a leopard at full speed to run through Pont des Arts?") print(result) ``` ### Creating Custom Tools ```python from smolagents import tool, CodeAgent, LiteLLMModel import os # Define a custom tool with the @tool decorator @tool def weather_forecast(city: str, days: int = 3) -> str: """Get weather forecast for a city. Args: city: The name of the city to get weather for days: Number of days to forecast (default: 3) Returns: A string with the weather forecast """ # In a real implementation, you would call a weather API here return f"Weather forecast for {city} for the next {days} days: Sunny, 25°C" # Create the agent with APIpie model model = LiteLLMModel( model_id="apipie/claude-3-opus-20240229", api_key=os.environ.get("APIPIE_API_KEY"), api_base="https://apipie.ai/v1" ) agent = CodeAgent( tools=[weather_forecast], model=model ) response = agent.run("What's the weather in Paris and should I pack an umbrella for my trip?") print(response) ``` ### Multi-Modal Agent with Vision ```python import os import base64 from pathlib import Path from smolagents import CodeAgent, LiteLLMModel, DuckDuckGoSearchTool # Configure APIpie with a vision-capable model model = LiteLLMModel( model_id="apipie/gpt-4o", # Use a vision-capable model api_key=os.environ.get("APIPIE_API_KEY"), api_base="https://apipie.ai/v1" ) # Create the agent agent = CodeAgent( tools=[DuckDuckGoSearchTool()], model=model ) # Read an image file def encode_image(image_path): with open(image_path, "rb") as image_file: return base64.b64encode(image_file.read()).decode('utf-8') # Create a query with an image image_path = "path/to/your/image.jpg" base64_image = encode_image(image_path) image_content = { "type": "image_url", "image_url": { "url": f"data:image/jpeg;base64,{base64_image}" } } # Run the agent with multimodal input response = agent.run([ {"role": "user", "content": [ {"type": "text", "text": "What's in this image? Can you identify the landmarks?"}, image_content ]} ]) print(response) ``` ### Sandboxed Execution with E2B ```python import os from smolagents import CodeAgent, LiteLLMModel, DuckDuckGoSearchTool from smolagents.code_interpreters import E2BInterpreter # Set up your E2B API key os.environ["E2B_API_KEY"] = "your-e2b-api-key" # Create APIpie model model = LiteLLMModel( model_id="apipie/mistral-large-2", api_key=os.environ.get("APIPIE_API_KEY"), api_base="https://apipie.ai/v1" ) # Create the agent with E2B sandboxed execution agent = CodeAgent( tools=[DuckDuckGoSearchTool()], model=model, code_interpreter=E2BInterpreter() # Use E2B for secure code execution ) response = agent.run("Analyze the GDP growth trends of the top 5 economies over the last decade and create a visualization.") print(response) ``` ### Sharing Tools to the Hub ```python from smolagents import tool, CodeAgent, LiteLLMModel import os @tool def currency_converter(amount: float, from_currency: str, to_currency: str) -> float: """Convert an amount from one currency to another. Args: amount: The amount to convert from_currency: The source currency code (e.g., USD, EUR) to_currency: The target currency code (e.g., USD, EUR) Returns: The converted amount """ # In a real implementation, you would call a currency API rates = {"USD": 1.0, "EUR": 0.92, "GBP": 0.78, "JPY": 153.2} if from_currency not in rates or to_currency not in rates: raise ValueError(f"Currency not supported: {from_currency} or {to_currency}") # Convert to USD first, then to target currency usd_amount = amount / rates[from_currency] target_amount = usd_amount * rates[to_currency] return round(target_amount, 2) # Share the tool to Hugging Face Hub currency_converter.push_to_hub("your-username/currency-converter-tool") ``` --- ## Agent CLI Tools Smolagents provides convenient CLI tools for quickly launching agents: ```bash # Run a general-purpose agent with web search capability smolagent "Plan a trip to Tokyo, Kyoto and Osaka between Mar 28 and Apr 7." \ --model-type "LiteLLMModel" \ --model-id "apipie/gpt-4o" \ --tools "web_search" # Run a web browsing agent webagent "Go to xyz.com/products, find the bestselling item, and summarize its features" \ --model-type "LiteLLMModel" \ --model-id "apipie/gpt-4o" ``` --- ## Troubleshooting & FAQ - **How do I ensure code execution is secure?**:br Use the E2B or Docker sandbox options to isolate agent code execution from your system. Never run untrusted code directly in your environment. - **Which APIpie models work best with Smolagents?**:br Models with strong code generation capabilities like GPT-4o, Claude 3 Opus, and DeepSeek Coder perform best for code agents. - **How do I handle rate limits?**:br You can implement retry logic in your agent setup or use LiteLLM's built-in retry mechanisms. - **Can I use multiple models in the same agent system?**:br Yes, each agent can use a different model. You might use a powerful model for reasoning and planning, and a more cost-effective model for simpler tasks. - **How do I debug agent execution?**:br Set `verbose=True` when creating your agent to see detailed logs of the agent's thinking and execution. - **What's the difference between CodeAgent and ToolCallingAgent?**:br CodeAgent writes its actions as Python code, which enables more complex logic and control flow. ToolCallingAgent uses a more traditional JSON-based tool calling approach. For more information, see the [Smolagents documentation](https://huggingface.co/docs/smolagents/){rel=""nofollow""} or the [GitHub repository](https://github.com/huggingface/smolagents){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} or [Hugging Face Discord](https://huggingface.co/join/discord){rel=""nofollow""} for assistance. # Vercel AI SDK Integration Guide ::div{.docs-image-row} ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"} ![Vercel AI SDK](https://apipie.ai/img/docs/integrations/Vercel-AI/vercel.png){height="125"} :: ::div{.docs-image-row} ![hero](https://apipie.ai/img/docs/integrations/Vercel-AI/hero.gif){height="125"} :: This guide will walk you through integrating the Vercel AI SDK with APIpie, enabling you to build powerful AI-powered applications with seamless streaming and tool calling capabilities. ## What is [Vercel](https://vercel.com/){rel=""nofollow""} AI SDK? Vercel AI SDK is a library designed to help developers build AI-powered user interfaces. It provides a set of tools and components for: - **Streaming text generation** from various language models - **Built-in React hooks** for chat interfaces and completions - **Tool calling capabilities** for executing functions based on AI requests - **Streaming UI patterns** for implementing typewriter effects and more - **Optimized performance** with edge runtime support By connecting Vercel AI SDK to APIpie, you unlock access to a wide range of powerful models while leveraging Vercel's optimized UI components and streaming capabilities. --- ## Integration Steps ### 1. Create an APIpie Account - **Register here:** [APIpie Registration](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Complete the sign-up process. ### 2. Add Credit - **Add Credit:** [APIpie Subscription](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Add credits to your account to enable API access. ### 3. Generate an API Key - **API Key Management:** [APIpie API Keys](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Create a new API key for use with Vercel AI SDK. ### 4. Install Vercel AI SDK Install the OpenAI-compatible provider for Vercel AI SDK: ```bash npm install @ai-sdk/openai-compatible # or yarn add @ai-sdk/openai-compatible # or pnpm add @ai-sdk/openai-compatible ``` ### 5. Configure Vercel AI SDK for APIpie Create a provider instance with your APIpie API key: ```typescript import { createOpenAICompatible } from '@ai-sdk/openai-compatible'; // Create a provider instance const provider = createOpenAICompatible({ name: 'apipie', apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); // Use the provider with a specific model const model = provider('gpt-4o-mini'); // or any model available on APIpie ``` --- ## Key Features - **Streaming Responses**: Get real-time tokens as they're generated for responsive UIs - **React Hooks**: Pre-built hooks for chat and completion interfaces - **Tool Calling**: Define and execute functions based on AI-driven decisions - **UI Components**: Build elegant AI interfaces with minimal code - **Edge Compatible**: Optimized for deployment on Vercel's Edge runtime --- ## Example Workflows | Application Type | What Vercel AI SDK Helps You Build | | ------------------------ | ----------------------------------------------------------- | | Chat Interfaces | Interactive conversational applications with streaming UIs | | Text Generation | Applications that generate and stream content to users | | AI Function Calling | AI agents that can make API calls, query databases, etc. | | Next.js AI Applications | Seamless integration of AI into your Next.js projects | | Multi-Modal Applications | Applications that handle both text and image inputs/outputs | --- ## Using Vercel AI SDK with APIpie ### Basic Text Generation ```typescript import { createOpenAICompatible } from '@ai-sdk/openai-compatible'; import { streamText } from 'ai'; export async function generateRecipe() { const apipie = createOpenAICompatible({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); const response = await streamText({ model: apipie('gpt-4o-mini'), prompt: 'Write a vegetarian lasagna recipe for 4 people.', }); // Process the streaming response for await (const chunk of response.textStream) { // Do something with each chunk as it arrives console.log(chunk); } // Or wait for the complete response await response.consumeStream(); return response.text; } ``` ### Tool Calling with Zod Schemas ```typescript import { createOpenAICompatible } from '@ai-sdk/openai-compatible'; import { streamText } from 'ai'; import { z } from 'zod'; export async function getWeatherWithAI() { const apipie = createOpenAICompatible({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); const response = await streamText({ model: apipie('gpt-4o-mini'), prompt: 'What is the weather in San Francisco, CA in Fahrenheit?', tools: { getCurrentWeather: { description: 'Get the current weather in a given location', parameters: z.object({ location: z.string().describe('The city and state, e.g. San Francisco, CA'), unit: z.enum(['celsius', 'fahrenheit']).optional(), }), execute: async ({ location, unit = 'celsius' }) => { // In a real application, this would call your weather API console.log(`Fetching weather for ${location} in ${unit}`); // Mock response return `The current weather in ${location} is 64°F.`; }, }, }, }); await response.consumeStream(); return response.text; } ``` ### Using with React Hooks ```tsx import { useChat } from 'ai/react'; import { createOpenAICompatible } from '@ai-sdk/openai-compatible'; // Create provider in a separate file export const apipie = createOpenAICompatible({ apiKey: process.env.NEXT_PUBLIC_APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); export default function ChatComponent() { const { messages, input, handleInputChange, handleSubmit } = useChat({ api: '/api/chat', // API route for server-side processing }); return (
{messages.map((m) => (
{m.content}
))}
); } ``` API route implementation (`/api/chat.js`): ```typescript import { createOpenAICompatible } from '@ai-sdk/openai-compatible'; import { streamText } from 'ai'; export const runtime = 'edge'; export async function POST(req) { const { messages } = await req.json(); const apipie = createOpenAICompatible({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); const response = await streamText({ model: apipie('gpt-4o-mini'), messages, }); return response.toTextStreamResponse(); } ``` --- ## Troubleshooting & FAQ - **Which models are supported?**:br Any model available via APIpie's OpenAI-compatible endpoint. - **How do I handle environment variables securely?**:br Store your API key in environment variables and never expose it in client-side code. For Next.js, use `.env.local` for development and Vercel environment variables for production. - **Can I use streaming responses with server components?**:br Yes, you can use streaming responses with React Server Components in Next.js App Router. - **How do I add custom headers to requests?**:br Use the `headers`option when creating your provider: ```typescript const provider = createOpenAICompatible({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', headers: { 'Custom-Header': 'value' }, }); ``` For more information, see the [Vercel AI SDK documentation](https://sdk.vercel.ai/docs){rel=""nofollow""} or the [GitHub repository](https://github.com/vercel/ai){rel=""nofollow""}. --- ## Support If you encounter any issues during the integration process, please reach out on [APIpie Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for assistance. # Seamless Migration from OpenAI: Step-by-Step Guide ![APIpie](https://apipie.ai/img/docs/apipie-logo.png){height="125" width="125"}![OpenAI](https://apipie.ai/img/docs/openai-white.png){height="108" width="400"} This guide outlines the simple steps required to migrate any application from [OpenAI](https://openai.com/){rel=""nofollow""} to APIpie, leveraging our compatible API structure. ## Introduction - Migrate from OpenAI If your application currently uses OpenAI, transitioning to APIpie is straightforward. Our API accepts the same structured requests, meaning you only need to update the base URL to migrate seamlessly. For more information on OpenAI's API structure, you might refer to [OpenAI API Reference](https://platform.openai.com/docs/api-reference){rel=""nofollow""}. ## Getting Started - Migration Steps ### 1. Create an Account - **Link**: [Register here](https://apipie.ai/profile/auth/register){rel=""nofollow""} - Follow the link and fill out the form to create your account. ### 2. Add Credit - **Link**: [Add Credit](https://apipie.ai/profile/subscribe){rel=""nofollow""} - Access the subscription section after logging in to add credits to your account. ### 3. Generate an API Key - **Link**: [Generate API Key](https://apipie.ai/profile/api-keys){rel=""nofollow""} - Navigate to the API keys section and create a new key. This key is necessary for API requests. ### 4. Update Your API Configuration - Locate where your application is configured to communicate with OpenAI, typically where the base URL for OpenAI is set. - Change the base URL from OpenAI's URL (e.g., `https://api.openai.com/v1`) to APIpie's URL `https://apipie.ai/v1`. - Update your configured API key to match the new APIpie API key ## Configuring OpenAI SDK to use APIpie APIpie's API endpoints for chat completions, vision, images, embeddings, and speech are fully compatible with OpenAI's API. If you have an application that uses one of OpenAI's libraries, you can quickly change it to point to APIpie, and start running your existing applications with our service. ### Python Configuration ```python import os import openai client = openai.OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), # Your APIpie API key base_url="https://apipie.ai/v1", # APIpie base URL ) ``` ### JavaScript/TypeScript Configuration ```typescript import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, // Your APIpie API key baseURL: 'https://apipie.ai/v1', // APIpie base URL }); ``` ## Querying Language Models ### Chat Completions Example ```python import os import openai client = openai.OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1", ) response = client.chat.completions.create( model="gpt-4o-mini", # APIpie supports various models including OpenAI ones messages=[ { role: 'system', content: 'You are an personal assistant' }, { role: 'user', content: 'Who won the 2015 NRL Grand Final?' }, ] ) print(response.choices[0].message.content) ``` ```typescript import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); const response = await client.chat.completions.create({ model: 'gpt-4o-mini', messages: [ { role: 'system', content: 'You are an personal assistant' }, { role: 'user', content: 'Who won the 2015 NRL Grand Final?' }, ], }); console.log(response.choices[0].message.content); ``` ## Streaming Responses ```python import os import openai client = openai.OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1", ) stream = client.chat.completions.create( model="gpt-3.5-turbo", messages=[ { role: 'system', content: 'You are an personal assistant' }, { role: 'user', content: 'Who won the 2015 NRL Grand Final?' }, ], stream=True, ) for chunk in stream: print(chunk.choices[0].delta.content or "", end="", flush=True) ``` ```typescript import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); async function run() { const stream = await client.chat.completions.create({ model: 'gpt-3.5-turbo', messages: [ { role: 'system', content: 'You are an personal assistant' }, { role: 'user', content: 'Who won the 2015 NRL Grand Final?' }, ], stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content || ''); } } run(); ``` ## Multimodal Vision Models ```python import os import openai client = openai.OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1", ) response = client.chat.completions.create( model="gpt-4-vision", messages=[{ "role": "user", "content": [ {"type": "text", "text": "What's in this image?"}, { "type": "image_url", "image_url": { "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg", }, }, ], }], ) print(response.choices[0].message.content) ``` ```typescript import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); const response = await client.chat.completions.create({ model: 'gpt-4.1-2025-04-14', messages: [ { role: 'user', content: [ { type: 'text', text: 'What is in this image?' }, { type: 'image_url', image_url: { url: 'https://en.wikipedia.org/wiki/Pie#/media/File:Arial_view_of_peach_pie_(722379748).jpg', }, }, ], }, ], }); console.log(response.choices[0].message.content); ``` Output: ```text This image captures a top-down view of a freshly baked peach pie, its golden-brown crust slightly uneven and beautifully rustic, hinting at a homemade charm. The filling is generous with thick, juicy slices of ripe peaches, their sunset-orange color peeking through the open gaps of the crust. The peaches glisten slightly, likely coated in a thin glaze of syrup or natural juices caramelized during baking. The pie's edges are rough and natural, not overly polished, giving it an inviting, cozy, farm-to-table feel. The background is plain and neutral, keeping full focus on the warm, delicious simplicity of the peach pie itself. ``` ## Image Generation ```python from openai import OpenAI import os client = OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1", ) prompt = """ A cheerful illustration of a fox and a rabbit painting a giant rainbow together in a sunny meadow. """ result = client.images.generate( model="dall-e-3", prompt=prompt ) print(result.data[0].url) ``` ```typescript import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); const prompt = ` A cheerful illustration of a fox and a rabbit painting a giant rainbow together in a sunny meadow. `; async function main() { const response = await client.images.generate({ model: 'dall-e-3', prompt: prompt, }); console.log(response.data[0].url); } main(); ``` ## Text-to-Speech ```python from openai import OpenAI import os client = OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1", ) speech_file_path = "speech.mp3" response = client.audio.speech.create( model="tts-1", input="Every great idea starts with a single step!", voice="alloy", ) response.stream_to_file(speech_file_path) ``` ```typescript import OpenAI from 'openai'; import * as fs from 'fs'; const client = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); async function main() { const speechFile = 'speech.mp3'; const mp3 = await client.audio.speech.create({ model: 'tts-1', voice: 'alloy', input: 'Every great idea starts with a single step!', }); const buffer = Buffer.from(await mp3.arrayBuffer()); await fs.promises.writeFile(speechFile, buffer); console.log(`Audio content written to ${speechFile}`); } main(); ``` ## Vector Embeddings ```python import os import openai client = openai.OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1", ) response = client.embeddings.create( model = "text-embedding-ada-002", input = "Sky is blue because air scatters sunlight’s blue wavelengths most." ) print(response.data[0].embedding) ``` ```typescript import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); const response = await client.embeddings.create({ model: 'text-embedding-3-large', input: 'Sky is blue because air scatters sunlight’s blue wavelengths most.', }); console.log(response.data[0].embedding); ``` Output ```text [0.8594738, 0.5930284, 0.90693754,... ] ``` ## Structured Outputs (JSON Mode) ```python from pydantic import BaseModel from openai import OpenAI import os, json client = OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1", ) class CalendarEvent(BaseModel): name: str date: str participants: list[str] completion = client.chat.completions.create( model="gpt-4o-mini", messages=[ {"role": "system", "content": "Extract the event information."}, {"role": "user", "content": "Tom and Collin are going to a concert on Saturday. Answer in JSON"}, ], response_format={ "type": "json_object", "schema": CalendarEvent.model_json_schema(), }, ) output = json.loads(completion.choices[0].message.content) print(json.dumps(output, indent=2)) ``` ```typescript import OpenAI from 'openai'; import { z } from 'zod'; const client = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); // Define schema with Zod const calendarEventSchema = z.object({ name: z.string(), date: z.string(), participants: z.array(z.string()), }); async function main() { const completion = await client.chat.completions.create({ model: 'gpt-4o-mini', messages: [ { role: 'system', content: 'Extract the event information.' }, { role: 'user', content: 'Tom and Collin are going to a concert on Saturday. Answer in JSON' }, ], response_format: { type: 'json_object', }, }); // Parse the result const output = JSON.parse(completion.choices[0].message.content); const validatedOutput = calendarEventSchema.parse(output); console.log(JSON.stringify(validatedOutput, null, 2)); } main(); ``` Output: ```text { "name": "Concert", "date": "Saturday", "participants": [ "Tom", "Collin" ] } ``` ## Tools (Functions Calling) ```python from openai import OpenAI import os, json client = OpenAI( api_key=os.environ.get("APIPIE_API_KEY"), base_url="https://apipie.ai/v1", ) tools = [{ "type": "function", "function": { "name": "get_weather", "description": "Get current temperature for a given location.", "parameters": { "type": "object", "properties": { "location": { "type": "string", "description": "City and country e.g. Bogotá, Colombia" } }, "required": [ "location" ], "additionalProperties": False }, "strict": True } }] completion = client.chat.completions.create( model="gpt-4o-mini", messages=[{"role": "user", "content": "What is the weather like in Dallas?"}], tools=tools, tool_choice="auto" ) print(json.dumps(completion.choices[0].message.model_dump()['tool_calls'], indent=2)) ``` ```typescript import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.APIPIE_API_KEY, baseURL: 'https://apipie.ai/v1', }); const tools = [ { type: 'function', function: { name: 'get_weather', description: 'Get current temperature for a given location.', parameters: { type: 'object', properties: { location: { type: 'string', description: 'City and country e.g. Bogotá, Colombia', }, }, required: ['location'], additionalProperties: false, }, strict: true, }, }, ]; async function main() { const completion = await client.chat.completions.create({ model: 'gpt-4o-mini', messages: [{ role: 'user', content: 'What is the weather like in Paris today?' }], tools, tool_choice: 'auto', }); console.log(JSON.stringify(completion.choices[0].message.tool_calls, null, 2)); } main(); ``` ## Benefits of Migrating to APIpie Migrating to APIpie provides several benefits: - **Full OpenAI SDK Compatibility**: Seamless transition with minimal code changes required. - **Multi-Provider Support**: Access to models from OpenAI, Anthropic, Google, Meta, and many more providers. - **Redundancy and Advanced Routing**: Never lose connectivity with redundant providers for all major models and advanced price or speed based routing. - **Cost Optimization**: Automatically route to the most cost-effective provider for your needs. - **Enhanced Performance**: Optimize for speed when needed with intelligent model routing. - **Simplified Billing**: One account, one API key, one invoice - regardless of how many AI providers you use. ## Need Help? If you encounter any issues during your migration or have further questions, reach out to us via [Discord](https://discord.gg/hs82THc9Tw){rel=""nofollow""}. # LLMs.txt > How to get AI tools like Cursor, Windsurf, GitHub Copilot, ChatGPT, and Claude to understand APIpie’s AI routing platform, providers, models, and APIs. ## What is LLMs.txt? **LLMs.txt** is a structured documentation format designed specifically for Large Language Models (LLMs). APIpie provides LLMs.txt files that expose **canonical, machine-readable context** about the APIpie platform so AI tools can reason correctly about: - What APIpie does - How AI routing works - Which providers and models are supported - How developers interact with the APIpie API These files are optimized for **AI consumption**, not humans. They contain structured summaries, references, and links to authoritative documentation. ## Available Routes APIpie publishes the following LLM context files: - **`/llms.txt`** A concise, high-signal overview of the APIpie platform (\~5–10K tokens). Intended for discovery, relevance checks, and standard context windows. - **`/llms-full.txt`** A comprehensive context file including routing concepts, provider behavior, API references, and deeper platform details (large context, 100K+ tokens). Both routes are public, cacheable, and stable. ## Choosing the Right File ::note **Start with `/llms.txt`.** It contains everything an AI tool needs to understand APIpie at a platform level and fits within standard LLM context limits. Use **`/llms-full.txt`** only if: - You need detailed API or routing behavior - Your AI tool supports large context windows (100K–200K+ tokens) - You want deeper implementation references :: ## What Information These Files Contain The APIpie LLM context files describe the platform at the **system level**, including: - Core purpose of APIpie (unified AI API, routing, aggregation) - Supported providers and model families - Routing, fallback, and observability concepts - High-level API structure and entry points - Links to authoritative documentation and MCP endpoints They do **not** include UI details, marketing language, or tutorials. ## Usage with AI Tools ### Cursor Cursor can use APIpie’s LLMs.txt files to provide better assistance when working with AI routing, provider selection, and API integration. #### How to use: 1. Reference the APIpie LLMs.txt URLs in chat 2. Add them to your project context using `@docs` Example: ```text @docs https://apipie.ai/llms.txt ``` [Read more about Cursor Web and Docs Search](https://docs.cursor.com/en/context/@-symbols/@-docs){rel=""nofollow""} ### Windsurf Windsurf can load APIpie’s LLM context files to understand routing behavior and platform APIs. #### Usage: - Reference `@docs` with the APIpie LLMs.txt URLs - Optionally create persistent workspace rules pointing to these files [Read more about Windsurf Web and Docs Search](https://docs.windsurf.com/windsurf/cascade/web-search){rel=""nofollow""} ### ChatGPT, Claude, and Other LLMs Any LLM that supports external documentation references can use APIpie’s LLM files. Examples: - “Using APIpie documentation from [https://apipie.ai/llms.txt”](https://apipie.ai/llms.txt%E2%80%9D){rel=""nofollow""} - “Follow the full APIpie routing and provider guidelines from [https://apipie.ai/llms-full.txt”](https://apipie.ai/llms-full.txt%E2%80%9D){rel=""nofollow""} These references help the model stay aligned with **current production behavior**, not training-time assumptions. ## Relationship to MCP LLMs.txt files and MCP serve different purposes: | Capability | LLMs.txt | MCP | | ---------- | ---------------------- | ------------------- | | Purpose | Platform understanding | Live interaction | | Access | Static, preloaded | Dynamic, tool-based | | Cost | Zero runtime cost | Per request | | Use case | Discovery, planning | Querying, execution | **LLMs.txt answers “what is APIpie and how does it work?”****MCP answers “interact with APIpie right now.”** Both are intentionally provided. ## Stability and Guarantees APIpie treats its LLMs.txt files as **public contracts**: - Routes are stable - Content reflects production behavior - Changes are intentional and reviewed - AI systems are expected to rely on them safely If an AI tool misunderstands APIpie, this is the **first place that gets fixed**. # MCP Server > Connect AI assistants (Cursor, Claude, VS Code, ChatGPT, Windsurf) to APIpie documentation using the Model Context Protocol (MCP). ## What This MCP Server Provides (Today) APIpie exposes a **read-only MCP server** that allows AI tools to: - **Discover documentation pages** - **Fetch page content** This MCP server is **documentation-only**. It does **not** provide live platform APIs, routing execution, provider metadata, pricing, or write access. ## MCP Endpoint Use this MCP server URL: ```text https://apipie.ai/mcp ``` ## Available Tools ### `list-pages` Lists all available documentation pages (title, path, description). | Parameter | Type | Description | | --------- | ----------------- | ---------------------- | | `locale` | string (optional) | Filter pages by locale | ### `get-page` Retrieves full markdown for a specific page. | Parameter | Type | Description | | --------- | ----------------- | ------------------------------------ | | `path` | string (required) | Page path (e.g. `/routing/overview`) | ## What This MCP Server Does NOT Do To avoid confusion for agent systems: - No model routing or execution - No provider/model inventory - No pricing/quotas - No account access - No mutations (create/update/delete) - No prompts (yet) If it’s not listed under **Available Tools**, it does not exist. --- # Connect an AI Client ## ChatGPT (Custom Connector) ::note Custom connectors are available in ChatGPT on the web for Plus/Pro plans, depending on workspace settings. :: 1. Open **ChatGPT Settings** 2. Go to **Connectors** 3. Enable **Developer mode** (if required) 4. Create a **New Connector** - **Name:** `APIpie Docs` - **MCP Server URL:** `https://apipie.ai/mcp` - **Authentication:** `None` 5. Save You can then enable the connector in a conversation and ask: - “List APIpie docs pages” - “Fetch the page for `/routing/overview` and summarize it” ## Claude Code ```bash claude mcp add --transport http apipie-docs https://apipie.ai/mcp ``` ## Claude Desktop 1. Open **Claude Desktop** → **Settings** → **Developer** 2. Click **Edit Config** 3. Add: ```json { "mcpServers": { "apipie-docs": { "command": "npx", "args": ["mcp-remote", "https://apipie.ai/mcp"] } } } ``` 4. Restart Claude Desktop ## Cursor ### Quick Install If your environment supports MCP deep links, add the server in Cursor via **Settings → Tools & MCP** using: - **Name:** `apipie-docs` - **Type:** `http` - **URL:** `https://apipie.ai/mcp` ### Manual Setup Create or edit: `.cursor/mcp.json` ```json { "mcpServers": { "apipie-docs": { "type": "http", "url": "https://apipie.ai/mcp" } } } ``` ## Visual Studio Code Create or edit: `.vscode/mcp.json` ```json { "servers": { "apipie-docs": { "type": "http", "url": "https://apipie.ai/mcp" } } } ``` > Ensure you have GitHub Copilot + Copilot Chat installed if you want Copilot to use it. ## Windsurf 1. Open **Windsurf** → **Settings** → **Cascade** 2. **Manage MCPs** → **View raw config** 3. Add: ```json { "mcpServers": { "apipie-docs": { "type": "http", "url": "https://apipie.ai/mcp" } } } ``` ## Zed Add to your Zed settings JSON: ```json { "context_servers": { "apipie-docs": { "source": "custom", "command": "npx", "args": ["mcp-remote", "https://apipie.ai/mcp"], "env": {} } } } ``` --- # Validate It Works After connecting, ask your assistant: - “Call `list-pages` and show me the first 10 results.” - “Call `get-page` for the routing overview page and summarize it.” If those work, the MCP connection is correct. --- # Usage Examples (Docs-Only) Once configured, AI assistants can reliably do: - “List all APIpie documentation pages” - “Get the authentication page and extract required headers” - “Find the docs page describing routing fallbacks” - “Retrieve the webhook docs and summarize event payloads” Anything involving live routing, models, or account state requires APIpie’s API (not MCP). # Features: Key Tools & Features for Developers ![Overview Feature Banner](https://apipie.ai/img/docs/features/overview-banner.svg) Welcome to the Features Overview for Neuronic AI's Apipie platform. We offer a cutting-edge solution for developers looking to leverage a wide variety of AI models across multiple providers, all under one API. Our platform aggregates leading AI services, including language models, image generation, voice synthesis, and embeddings, with support for hundreds of models like OpenAI, Anthropic, and ElevenLabs. Explore the innovative features that make us stand out in the competitive landscape: ## [Models Route](https://apipie.ai/docs/features/models){rel=""nofollow""} Easily retrieve and filter available AI models based on type, provider, and other criteria. The Models Route feature allows developers to efficiently select models tailored to their specific use cases. We provide one of the most comprehensive models route for AI services available on the market today. You will find much more than just a list of models to help you, with focus on knowledge developers need when developing or integrating with AI applications. ## [Completions](https://apipie.ai/docs/features/completions){rel=""nofollow""} Our Chat Completions endpoint provides a powerful yet flexible way to integrate conversational AI into your applications. This route is fully compatible with OpenAI's API structure, allowing for seamless integration. You can use it to build customer support bots, creative content generation tools, or interactive experiences. Whether using a single model or leveraging our multi-model query capabilities, this endpoint allows you to tap into the power of AI with ease. ## [Integrity](https://apipie.ai/docs/features/integrity){rel=""nofollow""} Our Integrity system minimizes AI hallucinations by querying models multiple times with an election process for selecting the most accurate response. This feature is ideal for businesses that require high accuracy in their AI outputs. ## [AI Model Pooling](https://apipie.ai/docs/features/pools){rel=""nofollow""} Enhance reliability and reduce data exposure with AI Model Pooling. This feature aggregates similar models into pools to ensure redundancy, higher rate limits, and failover mechanisms for consistent performance. Pools could have numerous renditions of the same model or dozens of models with like capabilities. Leverage our RAG tuning with pools and you will still get a knowledgeable response regardless the underlying models training. ## [Preferred Routing](https://apipie.ai/docs/features/routing){rel=""nofollow""} Choose between performance- or cost-optimized routing for your AI queries. Preferred Routing gives you control over how AI requests are handled across multiple providers. In the future, we will provide content-based routing that assesses the best model to use based on the content of your prompt. This will become increasingly important as different providers pay different data sources for training materials as the sector becomes more regulated. ## [Tools Support](https://apipie.ai/docs/features/tools){rel=""nofollow""} Configure tool usage across any model we offer. Choose between OpenAI or Anthropic models to handle your tool calls inline with any request you make to any model. Tools Calls are becoming increasingly popular, and more models are picking them up natively. We will continue to expand the options in tool call providers available on our platform. Tools are commonly used as a way for AI to request additional data from your application in a format your application can depend on. ## [RAG Tuning](https://apipie.ai/docs/features/ragtune){rel=""nofollow""} Leverage Retrieval-Augmented Generation (RAG) to enhance AI responses using your own data collections (one or more documents you upload to a collection). This feature provides a cost-effective alternative to training and fine-tuning models, and with larger context models, it offers similar value and accuracy with a much smaller barrier to entry. We have made it as easy as possible—upload documents to a collection, refer to the collection when you query any model we have, and it's that easy. ## [Pinecone Integration](https://apipie.ai/docs/features/pinecone){rel=""nofollow""} For developers who want or need more control over the vectorization process, we offer full Pinecone integration, allowing you to leverage Pinecone's vector database seamlessly in your AI workflows. This gives you full control over your vector storage and retrieval, optimizing your AI responses with custom embeddings. We will also be offering other vector and vector-alternative solutions in the future. ## [Image Generation](https://apipie.ai/docs/features/images){rel=""nofollow""} Access powerful image generation capabilities through our platform. Create, edit, and manipulate images using state-of-the-art AI models from various providers. Whether you need to generate original artwork, modify existing images, or create variations, our image generation feature provides a unified interface for all your visual AI needs. ## [Intelligent Model Management (IMM)](https://apipie.ai/docs/features/imm){rel=""nofollow""} Our Intelligent Model Management system optimizes your AI workflow by intelligently selecting and managing models based on your specific needs. IMM considers factors like performance, cost, and capabilities to ensure you're always using the most appropriate model for each task. ## [Voice Synthesis](https://apipie.ai/docs/features/Voices){rel=""nofollow""} Transform text into natural-sounding speech with our voice synthesis feature. Create lifelike voiceovers, generate audio content, and add voice capabilities to your applications using advanced text-to-speech models from leading providers. ## [Usage](https://apipie.ai/docs/features/usage){rel=""nofollow""} Track and analyze your API usage with unmatched transparency through our Usage feature. Apipie provides detailed insights into every API query, including token counts, response times, costs, and even source IPs. The **Historic Usage** API allows you to retrieve past query logs and audit performance, making it easier to optimize for cost-efficiency and application performance. No other provider offers this level of detailed query history accessible directly via API, empowering you to make data-driven decisions to maximize your AI integrations. Each of these features provides a unique advantage to developers and businesses looking to implement AI solutions in their applications. Whether you are focusing on accuracy, performance, or cost-efficiency, Neuronic AI offers a comprehensive platform to meet your needs. Explore the links to dive deeper into each feature. # Models Route Guide: Filter and Select AI Models ![Models Route Feature Banner](https://apipie.ai/img/docs/features/models-banner.svg){width="100%"} The Models API Route feature allows you to easily fetch and filter AI API Models available through our API, enabling precise selections based on model types, subtypes, providers, and more. This guide provides details on how to use the Models Route feature effectively for your AI applications. ## Fetching Models You can retrieve a list of available models or filter them based on specific criteria like type, subtype, and provider. Below is the API endpoint for fetching models: ```bash GET https://apipie.ai/v1/models ``` ### Query Parameters Here are the parameters you can use to filter the model results: - **type** (string): Filter by the type of model (e.g., `llm`, `vision`, `embedding`, `image`, `voice`, `moderation`, `coding`, `free`). :br Example: `llm` - **subtype** (string): Filter by the subtype of the model (e.g., `chat`, `fill-mask`, `question-answering`, `tts`, `stt`, `multimodal`). :br Example: `chat` - **provider** (string): Filter by the provider of the model (e.g., `openrouter`). :br Example: `openrouter` - **combination** (string): Combine filters for provider and subtype. :br Example: `provider=openrouter&subtype=chat` - **enabled** (integer): Filter only enabled models (`1` for enabled, `0` for disabled). :br Example: `enabled=1&subtype=chat` - **voices**: Retrieve a list of available voices for TTS models. :br Example: `voices` - **restrictions**: Retrieve a list of country restrictions for models. :br Example: `restrictions` ### Example API Requests Below are some example requests using curl: ```bash curl "https://apipie.ai/v1/models?type=llm" curl "https://apipie.ai/v1/models?subtype=multimodal" curl "https://apipie.ai/v1/models?type=embedding" curl "https://apipie.ai/v1/models?subtypetype=code" curl "https://apipie.ai/v1/models?type=llm&provider=openai" curl "https://apipie.ai/v1/models?type=llm&model=gpt-3.5-turbo-0125" ``` ### Pool Filtering: Subtype = Pool You can filter all available pools by setting the `subtype=pool` parameter in your API call. This lists all models grouped into pools. Example: ```bash curl "https://apipie.ai/v1/models?subtype=pool" ``` ### ChatX Filtering: Subtype = ChatX The `subtype=chatx` filter allows you to list all models that support streaming and memory services, ensuring a high-quality chat experience. These models are designed for real-time applications that require context retention across queries. Example: ```bash curl "https://apipie.ai/v1/models?subtype=chatx" ``` ### list voices: ?voices We will provde a list of all voices available from all our voice providers along with some voice details if available. ```bash curl -X GET 'https://apipie.ai/v1/models?voices' { "object": "list", "data": [ { "provider": "elevenlabs", "model": "eleven_multilingual_v2", "voice_id": "21m00Tcm4TlvDq8ikWAM", "name": "Rachel", "description": "description: calm, age: young, gender: female, accent: american, use case: narration" }, { "provider": "elevenlabs", "model": "eleven_multilingual_v1", "voice_id": "21m00Tcm4TlvDq8ikWAM", "name": "Rachel", "description": "description: calm, age: young, gender: female, accent: american, use case: narration" }, { "provider": "elevenlabs", "model": "eleven_monolingual_v1"... ``` ### list restrictions: ?restrictions We use this to list all the countries you are allowed to serve this AI to, serving any countries prohibited by any of these restricitons could cause significant consequences from APIpie to the underlyign aggregator and or the model provider themselves banning you and potentially with legal implications, you must not be a relay to work around these restrictions. If there are any providers not listed here, they are generally opensource and widely available without restriction. ```bash curl -X GET 'https://apipie.ai/v1/models?restrictions' { "object": "list", "data": [ { "restrict_id": 2, "name": "openai", "countries": "{\"countries\":[{\"name\":\"Albania\",\"alpha-2\":\"AL\"},{\"name\":\"Algeria\",\"alpha-2\":\"DZ\"},{\"name\":\"Afghanistan\",\"alpha-2\":\"AF\"},{\"name\":\"Andorra\",\"alpha-2\":\"AD\"},{\"name\":\"Angola\",\"alpha-2\":\"AO\"},{\"name\":\"Antigua and Barbuda\",\"alpha-2\":\"AG\"},{\"name\":\"Argentina\",\"alpha-2\":\"AR\"},{\"name\":\"Armenia\",\"alpha-2\":\"AM\"},{\"name\":\"Australia\",... ... ``` ## Model Record Structure Here’s a summary of the fields in each model record, if available: - `enabled`: Whether the model is currently available (1 = enabled, 0 = disabled). - `type`: The type of model (e.g., `llm`, `vision`, `embedding`). - `subtype`: The subtype of the model (e.g., `chat`, `tts`). - `provider`: The provider of the model (e.g., `openrouter`). - `route/model`: The API route for the specific model. - `description`: A description of the model's features and capabilities. - `max_tokens`: The maximum number of tokens for the model. - `max_response_tokens`: The maximum number of tokens in the model's response. - `req_price`: The cost of a token for the model's input. (Some providers dont tell us the price until after the query so we dont track them here) - `res_price`: The cost of a token for the model's output. - `price_type`: The type of pricing (e.g., per token, per request). - `price_unit`: The unit of pricing (e.g., USD). - `latency`: The latency at various prompt sizes, 2000 char/ 4000 char/ 8000 char/ 16000 char / 32000 char+ / Total average latency per query across all prompt sizes - `avg_cost`: The average cost of using the model for a request and response (can flcutuate based on usage for models that price input differently than output). - `query_count`: The number of queries made to the model. - `available`: Indicates whether the model is available for use. - `img`: The cost per image when using vision models and uploading images - `img_json`: The cost per image when generating images, various sizes and qualities are listed in json format ## Tips & Tricks - **Leverage Filters for Precision**: Use the `type`, `subtype`, and `provider` filters together to narrow down the exact models you need for specific tasks. For example, filtering by `subtype=chat` and `provider=openrouter` ensures you only receive chat-based models from OpenRouter. - **Use Pool and ChatX Subtypes**: The `subtype=pool` filter is ideal for querying models grouped into performance or use-case pools, while `subtype=chatx` ensures you're using models optimized for real-time chat applications with streaming and memory capabilities. - **Balance Cost and Capability**: Some models come with higher costs due to their advanced features (e.g., large token limits or memory support). Assess the requirements of your application and choose models that balance performance and cost. - **Use Voice and Restriction Filters**: If you're working with TTS or have specific geographic deployment needs, use `voices` or `restrictions` as query parameters to get more detailed options on voice models or country-specific availability. - **Test Different Providers**: Different providers may offer similar models with varying levels of accuracy, performance, and cost. Experiment with different providers to find the best match for your project. ## FAQs 1. **What is the primary use of the Models Route feature?** - The Models Route feature allows users to retrieve and filter available AI models based on type, subtype, provider, and more for targeted AI model integration. 2. **Can I filter models by both type and provider?** - Yes, you can combine filters such as `type=llm` and `provider=openai` to get a more refined list of models matching both criteria. 3. **What does the `subtype=chatx` filter do?** - The `chatx` subtype lists models that support streaming and memory services, providing a better experience for real-time chat applications. 4. **How can I retrieve available voices for TTS models?** - You can use the `voices` query parameter to fetch all available TTS voices. This will list models that support text-to-speech functionalities. 5. **Is the response time affected by the filters used?** - Response times are primarily affected by the specific model's latency and availability, rather than the filters applied. However, models with advanced features like memory or high token limits may introduce slightly longer response times. ## Links - [Fetch Models API Documentation](https://apipie.ai/docs/api/fetchmodels){rel=""nofollow""} ## Conclusion With this guide, you now have a clear understanding of how to use the Models Route feature to fetch and filter the AI models API effectively. Whether you need to query specific types, providers, or subtypes, our API offers robust capabilities to integrate the right models for your AI-powered projects. **Note:** Always review the cost implications based on the filters applied, as certain models may incur higher API usage fees due to their enhanced capabilities. # Chat Completions Guide: Multi-Provider Integrations ![Chat Completions Feature Banner](https://apipie.ai/img/docs/features/completions-banner.svg){width="100%"} Leverage the flexibility and power of our Chat Completions endpoint to integrate conversational AI into your applications. Whether you're building customer support bots, generating creative content, or crafting interactive experiences, our chat completion API is designed to help you get the most from AI models while providing advanced features like [RAG integration](https://apipie.ai/docs/features/ragtune), [model routing](https://apipie.ai/docs/features/routing), and [response verification](https://apipie.ai/docs/features/integrity). ## Chat Completions Overview The Chat Completions API allows developers to interact with AI models in a conversational manner by sending a sequence of messages to the model and receiving a response. The structure of these messages is fully compliant with OpenAI's API structure, allowing seamless integration. ### OpenAI Compatible Framework Our API is fully compatible with the [OpenAI Chat Completions API](https://platform.openai.com/docs/api-reference/chat){rel=""nofollow""}, making migration seamless. Just change the URL and your API key, and your OpenAI-based application can start using APIpie immediately. Once integrated, you can take advantage of our additional features: - [Automatic model routing](https://apipie.ai/docs/features/routing) for optimal performance and cost - [RAG integration](https://apipie.ai/docs/features/ragtune) for enhanced knowledge retrieval - [Response verification](https://apipie.ai/docs/features/integrity) to reduce hallucinations - [Usage tracking](https://apipie.ai/docs/features/usage) with detailed metrics - [Model pooling](https://apipie.ai/docs/features/pools) for high availability ## How Chat Completions Work A typical API call sends an array of messages to the chosen model, where each message consists of a role (either `system`, `user`, or `assistant`) and content. The model then processes the conversation and returns a response based on the provided context. ::note - We use the chat completions route for all LLM interfaces (instruct, text, etc.). All models are accessible through this unified interface. - For voice and image generation, see our dedicated [Images API](https://apipie.ai/docs/features/images) & [Voices API](https://apipie.ai/docs/features/voices) documentation. :: ### Available Model Providers Our API integrates with multiple leading AI providers, each offering unique capabilities: - [OpenAI](https://apipie.ai/docs/models/openai) - GPT-4, GPT-3.5 series - [Anthropic](https://apipie.ai/docs/models/claude) - Claude series - [Meta](https://apipie.ai/docs/models/llama) - Llama series - [Cohere](https://apipie.ai/docs/models/cohere) - Command series - [AI21](https://apipie.ai/docs/models/ai21) - Jurassic series - [Amazon](https://apipie.ai/docs/models/amazon) - Bedrock models - [Google](https://apipie.ai/docs/models/google) - Gemini series See our [Models Overview](https://apipie.ai/docs/features/models) for a complete list of supported models and providers. ### Example API Call Below is an example of how to use the Chat Completions API to generate a response: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "gpt-4o", "max_tokens": 100, "messages": [ { "role": "user", "content": "Why is the sky blue?" } ] }' ``` ### Response Example The expected response structure looks like the following: ```json { "id": "chatcmpl-5fde5f7fffe8d6dc1f18aab4a138d4b7", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "openai/gpt-4o", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "The sky appears blue primarily due to a phenomenon called Rayleigh scattering. Here's how it works:\n\n1. **Sunlight Composition**: Sunlight, or white light, is composed of many colors, each with different wavelengths. These colors can be seen in a rainbow or through a prism.\n\n2. **Atmospheric Interaction**: As sunlight reaches the Earth's atmosphere, it collides with molecules of gases and small particles.\n\n3. **Scattering and Wavelengths**: Different colors of light are" }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 13, "completion_tokens": 100, "total_tokens": 113, "prompt_characters": 20, "response_characters": 474, "cost": 0.001878, "latency_ms": 2727 }, "system_fingerprint": "fp_f4d98523ab4ae852" } ``` This structure provides the generated message, completion details, and detailed usage metrics for each request. ## Detailed Explanation of Parameters ### Required Parameters - **messages**(array): A sequence of message objects in the conversation. Each message has: - **role** (string): The role of the message in the conversation. Possible values: `system`, `user`, `assistant`. - **content** (string): The content of the message. ### Message Roles - **system**: This message type is used to guide the AI's behavior throughout the conversation. It sets the rules or parameters for how the AI should respond. For example, the system message might instruct the AI to speak in a specific language, adopt a particular tone, or adhere to certain boundaries in its replies. It's useful for defining the scope and role of the assistant, ensuring it stays within the desired context. :br**Example use case**: Defining the AI's behavior. ```json { "role": "system", "content": "You are a helpful assistant that speaks only in Swedish." } ``` - **user**: The user message represents the actual input or query from the person interacting with the AI. It's the prompt or question the user is submitting to the AI, and this message is crucial as it dictates what the AI will respond to. The content here is typically dynamic based on what the user is asking or requesting at any given moment. :br**Example use case**: The user's request to the AI. ```json { "role": "user", "content": "Why is the sky blue?" } ``` - **assistant**: This message type is the AI's response or any generated information that contributes to the conversation. It can be used for knowledge retrieval (such as in Retrieval-Augmented Generation, RAG) or historical memory when responding to ongoing conversations. The assistant message can be stored and used in context to improve the model’s ability to respond appropriately over a series of interactions. :br**Example use case**: AI's response or retrieved information. ```json { "role": "assistant", "content": "The sky is blue because of a phenomenon called Rayleigh scattering..." } ``` - **model** (string): Specifies the AI model to use. Normally a single base model name, you can provide up to 5 comma-separated models for multi-model queries which leverages our super query capabilities (Note: Some other features may not work with Super Query, such as pools for example) ### Optional Parameters - **provider** (string): Optionally, specify the AI provider, or omit this field to let the system choose the best-performing one. - **rag\_tune** (string): Use this to specify a RAG tune or vector collection for augmenting language model queries with data from a specified vector database. - **routing** (string): Defines how the call should be routed when multiple providers exist for a model. Options are: - `price`: Chooses the cheapest provider. - `perf`: Selects based on lowest latency for prompt size. - `perf_avg`: Chooses based on average latency. ### Memory Management Parameters Our [Integrated Model Memory (IMM)](https://apipie.ai/docs/features/imm) system provides advanced conversation context management that works across all supported models: - **memory** (integer): Enable memory management by setting to 1 - **mem\_session** (string): Unique identifier for separate conversation contexts - **mem\_expire** (integer): Set custom expiration time in minutes (max 1440, default 15) - **mem\_clear** (integer): Clear memory for a specific session when set to 1 Example API call with memory management: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "memory": 1, "mem_session": "user123", "mem_expire": 60, "messages": [ { "role": "user", "content": "Remember this: my favorite color is blue" } ] }' ``` This enables persistent conversation context across different models within the same session, allowing you to maintain context even when switching between different AI providers. ### Generation Control Parameters (Grouped) You can fine-tune the model's output by controlling randomness, creativity, and repetition through the following parameters: - **temperature** (number): Controls randomness in the output. A lower value makes the model more deterministic, while a higher value increases randomness. - **top\_p** (number): Cumulative probability cutoff for token selection. A lower value keeps the output more deterministic, while a higher value increases creativity. - **top\_k** (integer): Limits the token selection to the top-k most likely tokens. A lower value makes the response more predictable. - **frequency\_penalty** (number): Penalizes tokens based on their frequency in the output, promoting diversity. - **presence\_penalty** (number): Penalizes tokens that have already appeared, encouraging new content. Here’s a sample API call with these options: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "gpt-4o", "messages": [ { "role": "user", "content": "Tell me a story about a talking cat." } ], "temperature": 0.7, "top_p": 0.9, "top_k": 50, "frequency_penalty": 0.2, "presence_penalty": 0.5 }' ``` ### Beam Search Parameters Note: This is not available on all models, if the model you support does not support it, the parameter will be dropped. - **beam\_size** (integer): Determines the number of beams (sequences) to keep at each step during generation. This setting increases the probability of finding optimal sequences but requires more computational resources. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "llama-3.2-1b-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ], "beam_size": 5 }' ``` ### Other Parameters - **n** (integer): Number of completions to generate for a single input. - **max\_tokens** (integer): The maximum number of tokens to generate for each completion. - **stream** (boolean): If true, the API returns a stream of data chunks as the completion is generated, rather than waiting for the entire response. ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "gpt-4o", "messages": [ { "role": "user", "content": "How does streaming work in AI?" } ], "stream": true }' ``` ### Integrity and Verification Our unique [integrity feature](https://apipie.ai/docs/features/integrity) helps ensure response accuracy and reduce hallucinations: - **integrity** (integer): Controls the verification process where the model evaluates its own response. Set to 12 or 13 for different verification levels. - **integrity\_model** (string): Optionally specify a different model for verification checks. Learn more about response verification in our [Integrity Guide](https://apipie.ai/docs/features/integrity). ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openrouter", "model": "gpt-4o", "messages": [ { "role": "user", "content": "Why is the sky blue?" } ], "integrity": 13, "integrity_model": "gpt-3.5-turbo" }' ``` ### Example of Streaming Response ```json data: {"id":"chatcmpl-fc2e9121675cd68b07b6d1eb5e2b11e8","object":"chat.completion.chunk","created":1729535567,"provider":"openrouter","model":"openai/gpt-4o","choices":[{"index":0,"delta":{"role":"assistant","content":""},"logprobs":null,"finish_reason":null}],"system_fingerprint":"null"} data: {"id":"chatcmpl-fc2e9121675cd68b07b6d1eb5e2b11e8","object":"chat.completion.chunk","created":1729535567,"provider":"openrouter","model":"openai/gpt-4o","choices":[{"index":1,"delta":{"content":"The"},"logprobs":null,"finish_reason":null}],"system_fingerprint":"null"} ... data: {"id":"chatcmpl-fc2e9121675cd68b07b6d1eb5e2b11e8","object":"chat.completion.chunk","created":1729535567,"provider":"openrouter","model":"openai/gpt-4o","choices":[{"index":100,"delta":{"content":" sky"},"logprobs":null,"finish_reason":null}],"system_fingerprint":"null"} data: {"id":"chatcmpl-fc2e9121675cd68b07b6d1eb5e2b11e8","object":"chat.completion.chunk","created":1729535567,"provider":"openrouter","model":"openai/gpt-4o","choices":[{"index":101,"delta":{"content":""},"logprobs":null,"finish_reason":null}],"system_fingerprint":"null"} data: {"id":"chatcmpl-fc2e9121675cd68b07b6d1eb5e2b11e8","object":"chat.completion.chunk","created":1729535567,"provider":"openrouter","model":"openai/gpt-4o","choices":[{"index":102,"delta":{"content":""},"logprobs":null,"finish_reason":null}],"system_fingerprint":"null"} data: {"id":"chatcmpl-fc2e9121675cd68b07b6d1eb5e2b11e8","object":"chat.completion.chunk","created":1729535567,"provider":"openrouter","model":"openai/gpt-4o","choices":[{"index":0,"delta":{"content":""},"logprobs":null,"finish_reason":"stop"}],"system_fingerprint":"null"} data: {"id":"chatcmpl-fc2e9121675cd68b07b6d1eb5e2b11e8","object":"chat.completion.chunk","created":1729535567,"model":"openai/gpt-4o","system_fingerprint":"null","choices":[],"usage":{"prompt_tokens":13,"completion_tokens":100,"total_tokens":113,"prompt_characters":20,"response_characters":527,"cost":0.001878,"latency_ms":2831}} data: [DONE] ``` --- ## Usage Metrics in the Response We provide comprehensive usage data with every request to help you track costs and performance. Metrics include: - **prompt\_tokens**: Number of tokens in the input prompt. - **completion\_tokens**: Number of tokens generated in the model’s response. - **total\_tokens**: Sum of `prompt_tokens` and `completion_tokens`. - **prompt\_characters**: Number of characters in the input prompt. - **response\_characters**: Number of characters in the model’s response. - **cost**: The estimated cost of the request. - **latency\_ms**: Time taken for the model to generate a response. ### Example Response with Usage Metrics ```json { "id": "chatcmpl-5fde5f7fffe8d6dc1f18aab4a138d4b7", "object": "chat.completion", "created": 1729535643, "provider": "openrouter", "model": "openai/gpt-4o", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "The sky appears blue primarily due to a phenomenon called Rayleigh scattering. Here's how it works..." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 13, "completion_tokens": 100, "total_tokens": 113, "prompt_characters": 20, "response_characters": 474, "cost": 0.001878, "latency_ms": 2727 }, "system_fingerprint": "fp_f4d98523ab4ae852" } ``` **Note:** We are a leader in query usage reporting, offering extensive data tracking for every request. Additionally, historical billing data is available via API for audit purposes. ## Best Practices and Considerations ### Rate Limits and Quotas Each provider has specific rate limits and quotas. For optimal performance: - Use [model pools](https://apipie.ai/docs/features/pools) for high-volume applications - Implement proper error handling and retries - Monitor your [usage](https://apipie.ai/docs/features/usage) to stay within limits ### Model-Specific Considerations Different models have varying capabilities and limitations: - Check the [Models Overview](https://apipie.ai/docs/features/models) for specific model capabilities - Some models support longer context windows (up to 200K tokens) - Consider these factors when choosing a model: - Input/output token limits - Response latency requirements - Cost per token - Specialized capabilities (code, math, analysis) - Production requirements (SLA, availability) - Regulatory and compliance needs For detailed model comparisons and performance metrics, see our [Models Guide](https://apipie.ai/docs/features/models). ### Error Handling Common error scenarios and recommended handling: ```json { "error": { "message": "Error description", "type": "invalid_request_error", "param": "messages", "code": "context_length_exceeded" } } ``` - **context\_length\_exceeded**: Reduce input length or use a model with larger context window - **rate\_limit\_exceeded**: Implement exponential backoff or use [model pools](https://apipie.ai/docs/features/pools) - **invalid\_request\_error**: Check request parameters against the API specification ### Security Best Practices To ensure secure API usage: - Rotate API keys regularly - Use environment variables for key storage - Implement proper input validation - Consider using our [integrity checks](https://apipie.ai/docs/features/integrity) for sensitive applications # AI Image Generation: DALL-E, Stable Diffusion, & More ![Image Generation Feature Banner](https://apipie.ai/img/docs/features/images-banner.svg){width="100%"} Transform your text descriptions into stunning visuals with our powerful Image Generation API. Perfect for developers, designers, and businesses looking to automate image creation, generate marketing assets, or build creative applications. Our API supports multiple leading AI models and offers enterprise-grade reliability with detailed usage analytics. ## 🎨 Image Generation Overview The Image Generation API allows developers to create images from textual descriptions using state-of-the-art AI models. The API is fully compliant with OpenAI's API structure, enabling easy integration and migration from existing OpenAI-based applications. ### Open AI compatible framework Simply change the URL and your API key, and your application built on the OpenAI API can start using APIpie immediately. Once integrated, you can leverage all our additional features and capabilities. ## 🎯 Model Comparison and Capabilities The AI image generation landscape is rapidly evolving, with new models and capabilities being released regularly. Visit our [models route](https://apipie.ai/docs/features/models) for the most up-to-date list of supported image models and their capabilities. Here's a current comparison of major models: | Feature | DALL-E 3 | DALL-E 2 | Stable Diffusion XL | Stable Diffusion 2.1 | Flux | Amazon Titan | | ---------------------- | -------- | --------- | ------------------- | -------------------- | -------- | ------------ | | Style Control | ✅ | ❌ | ✅ | ✅ | ✅ | ✅ | | Image Editing | ❌ | ✅ | ✅ | ✅ | ✅ | ✅ | | Photorealistic Quality | High | Medium | High | Medium-High | High | High | | Processing Speed | Fast | Very Fast | Medium | Fast | Fast | Fast | | Cost | Higher | Medium | Lower | Lower | Medium | Medium | | Customization | Limited | Limited | Extensive | Extensive | Moderate | Moderate | ### Model Strengths and Use Cases **DALL-E 3** ([OpenAI Documentation](https://platform.openai.com/docs/guides/images){rel=""nofollow""}) - Photorealistic scenes and complex compositions - Accurate text rendering in images ([Research Paper](https://arxiv.org/abs/2304.14531){rel=""nofollow""}) - Detailed human faces and expressions - Architectural visualization Explore DALL-E 3's advanced capabilities in [OpenAI's Comprehensive Guide to DALL-E 3](https://openai.com/blog/dall-e-3){rel=""nofollow""} **DALL-E 2** ([Model Overview](https://openai.com/dall-e-2){rel=""nofollow""}) - Quick iterations and variations - Simple product images - Abstract art and patterns - Basic image editing and variations Understand DALL-E 2's features and limitations in the [Official DALL-E 2 System Card](https://github.com/openai/dalle-2-preview/blob/main/system-card.md){rel=""nofollow""} **Stable Diffusion Models** ([Stability AI](https://stability.ai/){rel=""nofollow""}) - SDXL: Enhanced photorealism and composition - SD 2.1: Balanced performance and quality ([GitHub Repository](https://github.com/Stability-AI/stablediffusion){rel=""nofollow""}) - Custom Models: Community-trained variations ([Hugging Face Hub](https://huggingface.co/models?pipeline_tag=text-to-image){rel=""nofollow""}) - Strengths: - Extensive customization options - Fine-tuning capabilities ([Guide](https://stability.ai/research){rel=""nofollow""}) - Model merging and custom training - Strong artistic style control - Open-source foundation ([License](https://github.com/Stability-AI/stablediffusion/blob/main/LICENSE){rel=""nofollow""}) **Flux** ([Official Documentation](https://blackforestlabs.ai/ultra-home/){rel=""nofollow""}) - Multiple model variants: - Pro models (FLUX.1-pro, FLUX.1.1-pro) for highest quality - Specialized models for depth and edge-guided generation - Performance models (schnell series) for faster results - Enterprise-grade reliability - Extensive customization options - Optimized for production use Explore Flux's technical details and implementation guides in the [Official Flux Model Repository](https://github.com/black-forest-labs/flux){rel=""nofollow""} **Amazon Titan** ([AWS Documentation](https://aws.amazon.com/bedrock/titan/){rel=""nofollow""}) - Enterprise-grade reliability - Consistent output quality - Good for production workloads Discover AWS Titan's enterprise capabilities in the AWS Titan Image Generation Guide ### Model Landscape and Evolution The image generation field is rapidly advancing, with new models and capabilities being released frequently. Key trends include: 1. **Increasing Quality** - Higher resolution outputs - Better photorealism - Improved text rendering - More accurate human features 2. **Enhanced Control** - Better style consistency - More precise prompting - Advanced editing capabilities - Fine-grained parameter control 3. **Customization Options** - Model fine-tuning - Custom training - Style adaptation - Domain specialization 4. **Performance Improvements** - Faster generation times - More efficient resource usage - Better scaling capabilities - Reduced costs For the most current information about available models and their capabilities, always refer to our [models documentation](https://apipie.ai/docs/features/models). We regularly update our supported models to include the latest advancements in the field. ## 🔄 How Image Generation Works A typical API call sends a text prompt describing the desired image, along with optional parameters to control the generation process. The API then processes this input and returns either a URL to the generated image or the image data directly as a base64-encoded string, depending on your specified response format. ### Example API Call Below is an example of how to use the Image Generation API to create an image: ```bash curl -L -X POST 'https://apipie.ai/v1/images/generations' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openai", "model": "dall-e-3", "prompt": "A serene lake at sunset with mountains in the background", "n": 1, "size": "1024x1024", "quality": "standard", "style": "natural" }' ``` ### Response Example The expected response structure looks like the following: ```json { "created": 1729535643, "data": [ { "url": "https://apipie.ai/temp/47g43g6f476vf.jpg", "revised_prompt": "A tranquil mountain lake reflecting the warm colors of sunset, with majestic peaks silhouetted against a golden sky, creating a peaceful and atmospheric natural scene" } ] } ``` ## 📝 API Parameters and Configuration ### Required Parameters - **prompt** (string): A text description of the desired image. Be specific and descriptive for best results. - **model** (string): Specifies the AI model to use for image generation (e.g., "dall-e-3", "dall-e-2", "stable-diffusion-3"). ### Optional Parameters - **provider** (string): Optionally specify the AI provider, or omit this field to let the system choose the best-performing one. - **n** (integer): Number of images to generate. Currently only supports generating 1 image at a time (n=1). - **size** (string): The size of the generated image. Available options: - "1024x1024" (default) - Supported by all models - "1024x1792" (portrait) - "1792x1024" (landscape) - "512x512" (legacy) - "256x256" (legacy) - **quality** (string): The quality of the generated image. Options: - "standard" (default) - "hd" (higher quality, may increase latency) - **style** (string): The stylistic approach for image generation. Options: - "natural" (default) - "vivid" (more vibrant and dramatic) - **response\_format** (string): The format of the response. Options: - "url" (default): Returns a URL to the generated image - "b64\_json": Returns the image as a base64-encoded string - **image** (string): Optional URL to an image you want the image model to update/modify. Note: Not supported by DALL-E 3. ### Example API Call with Optional Parameters ```bash curl -L -X POST 'https://apipie.ai/v1/images/generations' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "provider": "openai", "model": "dall-e-3", "prompt": "A futuristic cityscape at night with flying cars and neon lights", "n": 1, "size": "1024x1024", "quality": "hd", "style": "vivid", "response_format": "b64_json" }' ``` ### Response with Base64 Format ```json { "created": 1729535643, "data": [ { "b64_json": "iVBORw0KGgoAAAANSUhEUgAA...", "revised_prompt": "A sprawling futuristic metropolis illuminated by vibrant neon signs and streams of flying vehicles weaving between towering skyscrapers, creating a dynamic and energetic nighttime cityscape" } ] } ``` ## 💡 Common Use Cases Our Image Generation API supports a wide range of applications: 1. **E-commerce Product Visualization** - Generate product images from descriptions - Create lifestyle shots for marketing - Visualize custom product variations 2. **Content Creation** - Generate blog post illustrations - Create social media visuals - Design marketing materials 3. **Game Development** - Generate concept art - Create texture assets - Design character variations 4. **Educational Content** - Illustrate complex concepts - Create educational diagrams - Generate visual examples ## ⚡ Usage Metrics and Analytics We provide comprehensive usage data with every request to help you track costs and performance. Metrics include: - **cost**: The estimated cost of the request. - **latency\_ms**: Time taken for the image generation. - **image\_size**: Size of the generated image(s) in pixels. - **model\_used**: The specific model version used for generation. ### Example Response with Usage Metrics ```json { "created": 1729535643, "data": [ { "url": "https://apipie.ai/temp/47g43g6f476vf.jpg", "revised_prompt": "A serene lake at sunset with mountains in the background" } ], "usage": { "cost": 0.04, "latency_ms": 4521, "image_size": "1024x1024", "model_used": "dall-e-3" } } ``` **Note:** We provide detailed usage tracking for every request. Historical billing data is available via API for audit purposes. ## 🎓 Best Practices for Prompts ### Understanding Model Strengths and Specialties Different models excel at different types of images and use cases: **DALL-E 3** - Best for photorealistic images and complex scenes - Excellent at following detailed instructions - Strong at maintaining text accuracy in images - Ideal for commercial and professional use **DALL-E 2** - Excellent for artistic interpretations - Quick for rapid prototyping - Good for simple compositions - Efficient for basic image variations **Stable Diffusion Models** - SDXL: Superior composition and detail - SD 2.1: Excellent style control - Custom Models: Specialized for specific domains - Great for artistic and creative projects **Flux** - Pro variants: Best for high-quality commercial work - Schnell variants: Ideal for rapid prototyping - Specialized variants: Perfect for specific tasks (depth, edges) - Development variants: Great for testing and integration **Amazon Titan** - Reliable for production environments - Strong at maintaining brand guidelines - Good for batch processing - Consistent quality across variations ### Model-Specific Prompt Tips **DALL-E 3** - Be very specific about details - Use clear, structured descriptions - Specify camera angles and lighting - Include technical photography terms **Stable Diffusion** - Use artistic style keywords - Include composition keywords - Reference specific artists or movements - Experiment with negative prompts **Flux** - Balance descriptive and technical terms - Use clear style references - Include mood and atmosphere details - Specify brand guidelines when relevant **Amazon Titan** - Use professional terminology - Include technical specifications - Be precise about brand requirements - Structure prompts systematically ### Advanced Prompt Techniques To get the best results from the Image Generation API, consider these prompt engineering tips: 1. **Be Specific and Detailed** - Include style, lighting, perspective, and mood - Specify camera details for photographic results - Mention color schemes and textures 2. **Use Clear, Structured Language** - Avoid ambiguous descriptions - Structure prompts from general to specific - Use artistic terminology when relevant 3. **Consider Composition Elements** - Describe foreground and background - Specify focal points and subject placement - Include depth and perspective information 4. **Reference Artistic Styles** - Name specific art movements or artists - Describe technical aspects (brush strokes, medium) - Include time period references if relevant 5. **Follow Best Practices** - Ensure prompts comply with content guidelines - Test and iterate on prompts - Save successful prompt patterns ## ⚠️ Error Handling and Troubleshooting Common errors and their solutions: ### Rate Limiting ```json { "error": { "code": "rate_limit_exceeded", "message": "You have exceeded your rate limit. Please try again later." } } ``` Solution: Implement exponential backoff in your requests. ### Invalid Parameters ```json { "error": { "code": "invalid_parameter", "message": "The 'size' parameter must be one of: 1024x1024, 1024x1792, 1792x1024" } } ``` Solution: Verify parameter values against the documentation. ### Content Policy Violations ```json { "error": { "code": "content_policy_violation", "message": "Your request was rejected as it violates our content policy." } } ``` Solution: Review and adjust your prompt to comply with content guidelines. ### Performance Optimization Tips - Implement exponential backoff for rate limits - Use async requests for batch processing - Monitor and optimize request patterns - Consider using different models based on speed requirements ### Quality Improvement Tips - Try increasing the quality parameter to "hd" - Use more specific descriptions in your prompt - Consider using a different model better suited to your use case - Save successful prompts as templates - Break complex scenes into simpler components # RAG Tuning Guide: Enhance AI Responses ![RAG Tuning Feature Banner](https://apipie.ai/img/docs/features/ragtune-banner.svg){width="100%"} Unlock the full potential of AI with **RAG Tuning**—a revolutionary feature designed to simplify the process of tuning and optimizing AI responses using your custom data. By augmenting model queries with your own data, RAG Tuning ensures more accurate and context-aware responses, providing a cost-effective alternative to traditional training or fine-tuning. ## Why Use RAG Tuning? RAG Tuning allows you to bypass expensive and time-consuming model training by augmenting prompts with specific data from your own collection of documents. This enables any AI model to become highly specialized in responding to queries that relate to your content, making it ideal for businesses looking to integrate their own knowledge base without needing to retrain or fine-tune models. ## How RAG Tuning Works RAG Tuning is built around a few simple steps: 1. **Upload Your Document**: Upload any supported document type (PDF, DOC, DOCX, TXT, CSV, XLS, XLSX) to create a collection. 2. **RAG Tune Query**: Use the `rag_tune` parameter to specify the collection, enhancing the model's ability to provide answers based on your data. 3. **Tailored Responses**: The model retrieves relevant information from your documents and incorporates it into its response, providing answers that are informed by your data. ## Uploading a Document for RAG Tuning To start using RAG Tuning, upload a document and associate it with a collection. Here’s an API call that demonstrates how to do this: ```bash curl -L -X POST 'https://apipie.ai/ragtune' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: ' \ --data-raw '{ "collection": "my-ragtune-collection", "url": "https://example.com/mydocument.pdf", "metatag": "important-document" }' ``` This request processes the document and adds it to your specified collection. The `metatag` is optional but can be useful for categorizing your documents. ## Listing RAG Collections To see all collections you’ve created for RAG Tuning, you can use the following API request: ```bash curl -L -X POST 'https://apipie.ai/ragtune/listCollections' \ -H 'Accept: application/json' \ -H 'Authorization: ' ``` This call returns an array of collections currently associated with your account. ## Deleting a RAG Collection If you need to remove a collection from RAG Tuning, use this API call: ```bash curl -L -X POST 'https://apipie.ai/ragtune/deleteCollection' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: ' \ --data-raw '{ "collection": "my-ragtune-collection", "collectionName": "my-collection", "deleteAll": false, "ids": [ "id1", "id2" ], "filter": { "key": "value" } }' ``` This will delete specific documents from the collection based on the provided filters, or you can delete the entire collection if needed. ## Using RAG Tuning in Your Query Once you’ve uploaded documents and created a collection, you can augment your queries with RAG Tuning. Here’s an example of how to use the `rag_tune` parameter in an API query: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "messages": [ { "role": "user", "content": "Why the sky is blue?" } ], "model": "gpt-3.5-turbo", "provider": "openai", "rag_tune": "my-ragtune-collection" }' ``` In this example, the `rag_tune` parameter ensures that the query pulls relevant data from the specified RAG collection to generate a more accurate response. ## Benefits of RAG Tuning for Businesses - **Cost-Effective**: Unlike traditional model training, RAG Tuning is less expensive, as it involves augmenting prompts rather than retraining the model. - **Flexible**: Use any supported model for RAG Tuning, including OpenAI models, while incorporating your own data into the responses. - **Efficient**: The process is quick and straightforward, making it accessible to businesses looking for tailored AI responses without the complexity of fine-tuning. - **Simple Vector-Based Approach**: Our RAG Tuning simplifies the use of vector-based search, providing a more user-friendly alternative to direct vector databases like Pinecone. ## RAG Tuning Fees Keep in mind that RAG Tuning incurs the following costs: - **Document Upload**: Each document uploaded to a RAG collection is charged. - **Querying Documents**: Additional fees apply when fetching relevant data from the uploaded documents. - **Storage**: Daily fees apply for storing documents and maintaining collections. ## Tips & Tricks - **Selective Use of Collections**: Only apply RAG Tuning to specific queries where augmenting the model with your data is beneficial. - **Experiment with Models**: While `gpt-4o` is a great default, experiment with different models to find the most cost-effective solution for your needs. - **Manage Costs**: Keep track of the number of documents and collections to balance costs with the accuracy gains from RAG Tuning. ## FAQs 1. **How do I create an index for RAG Tuning?** - You don’t need to manually create an index or namespace. When you upload a document and assign it to a `collectionName`, it automatically handles both the index and namespace for you. 2. **Can I use RAG Tuning with any AI model?** - Yes, RAG Tuning works with any AI model we support, allowing you to augment your queries with your own data regardless of the model used. 3. **Is there a limit to the number of documents I can upload to a collection?** - There is no hard limit, but additional charges apply based on the number of documents uploaded and stored. Monitor your usage to manage costs effectively. 4. **How does RAG Tuning affect the model’s response time?** - RAG Tuning may increase response times slightly as the system fetches relevant data from your documents to augment the model’s answer. However, this trade-off usually results in more accurate and context-aware responses. 5. **What happens to my data after I delete a collection?** - Once a collection is deleted, all associated documents and embeddings are permanently removed from the system, ensuring your data remains secure and private. ## Links - [RAG Tuning API Documentation](https://apipie.ai/docs/api/ragtune){rel=""nofollow""} - [Chat Completions API](https://apipie.ai/docs/api/chatcompletions){rel=""nofollow""} ## Conclusion RAG Tuning empowers businesses to enhance their AI models without the cost and complexity of retraining. By augmenting model queries with your own data, RAG Tuning enables any model to provide more accurate and contextually relevant answers. Explore the power of RAG Tuning today and make your AI integrations smarter, faster, and more reliable. # AI Voice: ElevenLabs & OpenAI TTS API Guide ![Voice Generation Feature Banner](https://apipie.ai/img/docs/features/voices-banner.svg){width="100%"} Transform text into natural-sounding speech with our enterprise-ready Voice Synthesis API. Perfect for developers, content creators, and businesses looking to add professional voice capabilities to their applications. Our API supports leading AI models from [ElevenLabs](https://elevenlabs.io){rel=""nofollow""} and [OpenAI](https://platform.openai.com/docs/guides/text-to-speech){rel=""nofollow""}, offering 30+ premium voices, multilingual support, and advanced customization options with enterprise-grade reliability. ::note For detailed information about using models with APIpie, check out our [Models Overview](https://apipie.ai/docs/features/models) and [Completions Guide](https://apipie.ai/docs/features/completions). :: ## 🎙️ Voice Synthesis Overview The Voice Synthesis API allows developers to convert text into high-quality speech using state-of-the-art AI models. The API supports both [ElevenLabs](https://elevenlabs.io){rel=""nofollow""} and [OpenAI](https://platform.openai.com/docs/guides/text-to-speech){rel=""nofollow""} voice models, providing: - 30+ premium voices across multiple languages and accents - Advanced voice customization and style control - Real-time voice generation with low latency - Enterprise-grade reliability and scalability - Comprehensive usage analytics and monitoring - Secure API access with rate limiting ## 🎯 Model Comparison and Capabilities ### Model Feature Comparison | Feature | ElevenLabs | OpenAI TTS | | -------------------- | ---------------------------------- | --------------------------- | | Voice Quality | High fidelity with emotion control | Professional studio quality | | Language Support | 30+ languages | Primary focus on English | | Generation Speed | Variable (Flash to Standard) | Consistently fast | | Customization | Extensive voice settings | Basic voice selection | | Cost Efficiency | Pay per character | Pay per character | | Real-time Generation | Yes (with Flash models) | Yes | | Voice Cloning | Available | Not available | | Enterprise Support | Yes | Yes | ### ElevenLabs Models ::note Visit our [Models Overview](https://apipie.ai/docs/features/models) for the most up-to-date list of supported voice models and their capabilities. :: | Model | Description | Max Tokens | Provider | | ------------------------ | ----------------------------------------------- | ---------- | ---------- | | eleven\_multilingual\_v2 | Latest multilingual model with enhanced quality | 5000 | elevenlabs | | eleven\_multilingual\_v1 | First generation multilingual model | 5000 | elevenlabs | | eleven\_monolingual\_v1 | English-optimized model | 5000 | elevenlabs | | eleven\_turbo\_v2 | Fast generation model | 5000 | elevenlabs | | eleven\_turbo\_v2\_5 | Enhanced turbo model | 5000 | elevenlabs | | eleven\_flash\_v2 | Ultra-fast generation | 5000 | elevenlabs | | eleven\_flash\_v2\_5 | Latest ultra-fast model | 5000 | elevenlabs | ### OpenAI Models | Model | Description | Provider | | ---------- | ---------------------------- | -------- | | tts-1-hd | High-definition voice models | openai | | tts-1-1106 | Standard voice models | openai | ## 🗣️ Available Voices ### Technical Specifications | Specification | Details | | ---------------- | ------------------ | | Audio Format | MP3, WAV | | Sample Rate | 16kHz - 48kHz | | Bit Depth | 16-bit, 24-bit | | Channels | Mono, Stereo | | Latency | 200ms - 2000ms | | Max Input Length | 5000 tokens | | Rate Limiting | Yes (configurable) | ### ElevenLabs Voices \::note{to="[https://elevenlabs.io/voice-library">}](https://elevenlabs.io/voice-library%22%3E%7D){rel=""nofollow""} Browse more voices in the ElevenLabs Voice Library \:: #### Professional Narration - **Rachel**: Young female, American accent, calm tone - ideal for narration - **Drew**: Middle-aged male, American accent - perfect for news reading - **Antoni**: Young male, American accent - well-rounded narrator - **Thomas**: Young male, American accent - calm meditation voice - **Bill**: Older male, American accent - trustworthy narration #### Character Voices - **Clyde**: Middle-aged male, American accent - war veteran character - **Dave**: Young male, British-Essex accent - conversational gaming voice - **Fin**: Older male, Irish accent - sailor character - **Glinda**: Middle-aged female, American accent - witch character - **Charlotte**: Young female, Swedish accent - seductive character #### News & Media - **Paul**: Middle-aged male, American accent - ground reporter - **Sarah**: Young female, American accent - soft news voice - **Daniel**: Middle-aged male, British accent - authoritative news - **Alice**: Middle-aged female, British accent - confident news - **Joseph**: Middle-aged male, British accent - field reporter ### OpenAI Voices #### HD Voices - **Shimmer**: Clear and expressive - **Alloy**: Versatile and balanced - **Echo**: Warm and natural - **Fable**: Engaging storyteller - **Onyx**: Deep and authoritative - **Nova**: Bright and energetic ## 📝 API Parameters and Configuration ::note For detailed API documentation and integration guides, visit our [API Reference](https://apipie.ai/docs/category/api/audio). :: ### Required Parameters - **model**: The AI model to use for voice generation - **voice**: The specific voice to use - **input**: The text to convert to speech ### Optional Parameters (ElevenLabs) - **stability** (0-1): Controls voice stability - **similarity\_boost** (0-1): Enhances similarity to the original voice - **style** (0-1): Adjusts speaking style intensity - **use\_speaker\_boost** (boolean): Enhances speaker clarity ## 💡 Example API Calls ### ElevenLabs Example ```bash curl -X POST 'https://apipie.ai/v1/audio/speech' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer YOUR_API_KEY' \ --data-raw '{ "model": "eleven_multilingual_v2", "voice": "Rachel", "input": "Hello! This is a test of the ElevenLabs text to speech API.", "voice_settings": { "stability": 0.5, "similarity_boost": 0.75 } }' ``` ### OpenAI Example ```bash curl -X POST 'https://apipie.ai/v1/audio/speech' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer YOUR_API_KEY' \ --data-raw '{ "model": "tts-1-hd", "voice": "shimmer", "input": "Hello! This is a test of the OpenAI text to speech API." }' ``` ## 📊 Response Examples ### ElevenLabs Response ```json { "created": 1729535643, "audio": { "content_type": "audio/mpeg", "url": "https://example.com/generated-audio.mp3" }, "usage": { "text_characters": 57, "cost": 0.004275, "latency_ms": 1200 } } ``` ### OpenAI Response ```json { "created": 1729535643, "audio": { "content_type": "audio/mpeg", "url": "https://example.com/generated-audio.mp3" }, "usage": { "text_characters": 52, "cost": 0.0035, "latency_ms": 800 } } ``` ## 🎯 Common Use Cases 1. **Content Creation** - Audiobook production - Podcast generation - Video narration - E-learning content 2. **Entertainment** - Game character voices - Animation dubbing - Interactive storytelling - Voice-enabled NPCs 3. **Business Applications** - IVR systems - Virtual assistants - Customer service - Corporate training 4. **Accessibility** - Screen readers - Text-to-speech for visually impaired - Language learning tools - Reading assistance ## ⚡ Best Practices 1. **Model Selection** - Use multilingual models for multiple language support - Use turbo/flash models for faster generation - Use HD models for highest quality output 2. **Voice Selection** - Choose voices based on use case - Consider accent and age appropriate for content - Test multiple voices to find the best fit 3. **Text Preparation** - Use punctuation to control pacing - Break long text into natural segments - Include phonetic spelling for unusual words 4. **Performance Optimization** - Cache frequently used audio - Implement proper error handling - Monitor usage and costs ## ⚠️ Error Handling Common errors and solutions: ```json { "error": { "code": "invalid_voice", "message": "The specified voice is not available for this model." } } ``` Solution: Verify voice compatibility with chosen model. ```json { "error": { "code": "text_too_long", "message": "Input text exceeds maximum length for selected model." } } ``` Solution: Break text into smaller segments. ## 🔒 Security and Ethics - Voice generation requires responsible use - Implement appropriate content filtering - Monitor for potential misuse - Secure API access and authentication - Respect voice rights and permissions ## 📚 Additional Resources - [ElevenLabs Documentation](https://docs.elevenlabs.io){rel=""nofollow""} - [OpenAI TTS Documentation](https://platform.openai.com/docs/guides/text-to-speech){rel=""nofollow""} - [ElevenLabs Voice Library](https://elevenlabs.io/voice-library){rel=""nofollow""} - [ElevenLabs Blog](https://elevenlabs.io/blog){rel=""nofollow""} - [OpenAI TTS Examples](https://platform.openai.com/examples?category=speech){rel=""nofollow""} - [Speech Synthesis Technology Overview](https://arxiv.org/abs/2106.15561){rel=""nofollow""} # Configuration State Management ![State Management Feature Banner](https://apipie.ai/img/docs/features/state/state-mgmt.svg){width="100%"} APIpie by Neuronic AI introduces the **first-of-its-kind State Management system** for AI applications — an innovation you won’t find anywhere else. Built directly into our platform, this feature empowers developers, enterprises, and even open-source projects to take full control of AI behavior at runtime, without rewriting a single line of code. With APIpie State Management, every configuration — from **model selection** and **memory settings** to **integrity controls**, **search integration**, and even the **system prompt** — can be centrally managed, persisted, and enforced. Whether you’re sending queries through our **OpenAI-compliant API**, shaping behavior in real time with our **Inline CLI**, or governing application defaults through a **visual dashboard**, state becomes the backbone of flexible and scalable AI deployment. --- The value is enormous: - **For developers**, it’s a flexible **configuration abstraction layer**. Just like Kubernetes abstracted infrastructure, APIpie abstracts AI configuration. Your apps and SDKs no longer need to bake in rigid settings or defaults. Instead, state persists across requests, overriding whatever the app enforces — giving you a powerful layer to manage models, memory, integrity, and search behavior without constant rewrites. - **For enterprises**, it provides a **single point of governance** across teams and applications. Policies can be updated globally, costs optimized, and compliance enforced — all in real time. And when needed, state changes can be locked down so they’re only honored through the GUI, ensuring strict oversight while still enabling rapid adaptation. - **For AI enthusiasts**, it’s the ultimate power-up. Take any open-source AI app — even one hard-coded to use GPT — and instantly make it run with Claude, Gemini, or any other model. Add memory where none existed, change behavior on the fly, and unlock capabilities the original developers never built in. State Management gives you control over tools you already love, without touching a single line of code. - **For innovators**, it’s a launchpad. By moving configuration outside of applications, APIpie opens the door to entirely new workflows, integrations, and value-adds. Features can be layered, swapped, or extended without changing code — creating possibilities we’ve only begun to explore. --- This is not just convenience — it’s competitive advantage. With APIpie, businesses and builders can adapt faster, experiment freely, and push AI-driven products to market with confidence. Our State Management framework is the first of many **industry firsts** we’re bringing to the table, giving you an edge over competitors and empowering you to harness AI more effectively than ever before. **three powerful ways to manage state** across applications, API keys, and users. Whether you need **real-time control inside a prompt**, **programmatic management via API**, or a **visual interface**, our platform ensures full flexibility. - **GUI** – Manage app or API key state visually via the APIpie dashboard - **[API](https://apipie.ai/docs/api/get-current-state-settings)** – Use the `/v1/state` endpoint to programmatically create, update, or delete state for apps and or users. - **[Inline CLI](https://apipie.ai/docs/features/inlinecli)** – Manage state directly in your prompts with natural commands like `:setmodel:`. --- ## GUI State Management The APIpie **State Management GUI** provides a visual interface for configuring and controlling state across apps and API keys. It’s designed for enterprises, teams, and developers who prefer a centralized, intuitive way to enforce AI settings without relying solely on API calls or inline CLI commands. Each tab in the GUI corresponds to a configuration area — from general flags and model selection to fine-grained controls, memory management, and internet search behavior. On the right-hand side of the GUI, you’ll always see the **Current Settings JSON**, which shows the effective configuration that will override any incoming API call settings for this app or user scope. API keys can be managed through the [API-key interface](https://apipie.ai/profile/api-keys){rel=""nofollow""}:br Click on "Manage" next to the API key you want to manage --- ### General Settings ![State Management - General Settings](https://apipie.ai/img/docs/features/state/state-general.png){width="100%"} The **General tab** provides high-level switches that control scope, governance, and CLI access. - **enable\_user\_states** – Enables per-user scoped state. When `true`, every unique `user` value passed with this API key has its own isolated state. When `false`, the API key operates in **app-scoped** mode. - **enable\_inline\_cli** – Allows users to adjust state dynamically inside prompts with [Inline CLI](https://apipie.ai/docs/features/inlinecli). Disable this to lock down state so prompts can’t alter configuration. - **gui\_only** – Locks configuration changes to the GUI only. When enabled, state cannot be modified via API or Inline CLI. Ideal for enterprises enforcing strict governance. - **user** – Used to force all completions through a specific `user` value when calling APIs with this key (not compatible with user-scoped management). --- ### Models ![State Management - Model Settings](https://apipie.ai/img/docs/features/state/state-models.png){width="100%"} The **Models tab** allows you to define which AI model to use as the default for this state. - **Provider** – Choose from supported AI providers (e.g., OpenAI, Anthropic, Google). - **Model** – Select a specific model from the chosen provider. For example: `openai/gpt-5-chat-latest`. - **Set default model** – When enabled, forces all completions under this key or user to use the chosen model unless overridden by Inline CLI (if allowed). --- ### Controls ![State Management - Controls Settings](https://apipie.ai/img/docs/features/state/state-controls.png){width="100%"} The **Controls tab** lets you fine-tune model generation parameters for creativity, coherence, and randomness. - **Temperature (0.0 → 2.0)** – Controls randomness. Lower = more deterministic, higher = more creative. - **Top-p (0.0 → 1.0)** – Nucleus sampling cutoff. Keeps responses focused while balancing diversity. - **Top-k (0 → 500+)** – Restricts token selection to the top-k most probable. Useful for reducing noise. - **Frequency Penalty (-2.0 → +2.0)** – Reduces repetition of frequently used tokens. - **Presence Penalty (-2.0 → +2.0)** – Discourages tokens already present in the conversation to encourage novelty. --- ### Memory ![State Management - Memory Settings](https://apipie.ai/img/docs/features/state/state-memory.png){width="100%"} The **Memory tab** manages both short-term and long-term AI memory. - **memory** – Global toggle for enabling memory in this state. - **mem\_session** – Optional custom identifier for separate memory chains (e.g., `customer-123-sessionA`). - **mem\_expire (min)** – Sets expiration in minutes for stored memory (time-based). - **Short-term memory pairs** – Number of recent message pairs (Q\&A) retained in memory (max 10). - **Long-term memory pairs** – Number of older message entries recalled from long-term memory (max 10). This enables features like personalized conversations, tenant-specific persistence, or multi-session separation. --- ### Internet ![State Management - Internet Settings](https://apipie.ai/img/docs/features/state/state-internet.png){width="100%"} The **Internet tab** configures web search augmentation and content filters. - **search\_whitelist** – Comma-delimited list of domains explicitly allowed for AI search (e.g., `apnews.com, reuters.com`). - **search\_blacklist** – Comma-delimited list of domains blocked from AI search (e.g., `cnn.com, foxnews.com`). - **Search Depth (low, medium, high)** – Determines how much web content/context is pulled into completions. - **search\_lang** – Restrict search results to a specific language (ISO-2 code). - **search\_geo** – Restrict search results by geographic region (ISO-2 country code). These settings allow enterprises to enforce compliance, control bias in sources, and optimize retrieval quality. --- Together, these **GUI tabs** give teams complete control over every aspect of AI behavior, while the **Current Settings JSON** ensures transparency by showing exactly what configuration is active. Whether you want flexibility for developers or governance for enterprises, the GUI provides a straightforward and powerful state management experience. --- ## API State Management Developers can manage application or user state programmatically via the `/v1/state` endpoint. This provides fine-grained control over configuration, memory, and routing at both **per-app** and **per-user** levels. **[State Management API REFERENCE](https://apipie.ai/docs/api/get-current-state-settings)** ### Authentication Authenticate using either: - **Authorization header** with a Bearer token - **x-api-key header** ```http Authorization: Bearer x-api-key: ``` ### Endpoints #### **GET /v1/state** Retrieve the current state settings. - Without query params → returns app-level state. - With `?user={userId}` → returns user-specific state. - With `?app_name={name}` → supports centralized enterprise state management. ```bash curl -X GET "https://apipie.ai/v1/state" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` #### **POST /v1/state** Create or update state. Supports **partial updates**, **key deletions**, and **toggling features**. ```bash curl -X POST "https://apipie.ai/v1/state" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "settings": { "memory": false, "model": "openai/gpt-4-1", "search_lang": "en", "search_geo": "us", "shortMem": 10, "routing": "price" } }' ``` #### **DELETE /v1/state** Delete entire state records. - With `?user={userId}` → deletes state for a user. - Without query → deletes app-level state. ```bash curl -X DELETE "https://apipie.ai/v1/state" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` --- ## Example API Workflows ### Disable Inline CLI ```bash curl -X POST "https://apipie.ai/v1/state" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "enable_inline_cli": false, "settings": {} }' ``` ### Add & Delete State Keys in One Request ```bash curl -X POST "https://apipie.ai/v1/state" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "settings": { "nlp": false }, "delete": ["routing"] }' ``` ### Enable User Scoped States ```bash curl -X POST "https://apipie.ai/v1/state" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"enable_user_states": true }' ``` --- ## Example Response All state responses return a consistent schema including scope, key, settings, and feature flags: ```json { "scope": "app", "key": "stateTest", "settings": { "memory": false, "mem_session": "testState", "model": "openai/gpt-4-1", "search_lang": "en", "search_geo": "us", "shortMem": 10, "longMem": 2, "routing": "price" }, "enable_inline_cli": true, "gui_only": false } ``` --- ## Inline CLI State Management The [Inline CLI](https://apipie.ai/docs/features/inlinecli) provides **real-time state control** inside your prompt flow. This feature lets you configure and persist AI state with natural commands in the first or last 250 characters of any prompt. ### Quick Start Send `:help` in your prompt with your APIpie key in your favorite chat app to see all supported commands. This is the fastest way to explore available options. ### Example Usage The following example shows how to **persistently set a model** and then confirm the saved state with `:getstate`: ```bash curl -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Authorization: Bearer ' \ -H 'Content-Type: application/json' \ --data-raw '{ "user": "12345", "inline_cli": "all", "messages": [ { "role": "user", "content": ":setmodel:openai/gpt-4o :getstate" } ] }' ``` This sets the model to **OpenAI GPT-4o** for all subsequent requests (until unset) and immediately returns all current state configuration for this user or app depending on if its scoped for app or user. Youc an do a lot more than jsut state management with our inline CLI, for example: Inline CLI supports commands for: - **Viewing state** (`:getstate`) - **Model selection** (`:setmodel:openai/gpt-4o`, `:unsetmodel`) - **Behavior shaping** (`:becreative`, `:beprecise`) - **Integrity controls** (`:setintegrity`, `:answersuperintegrity`) - **Memory control** (`:setmemoryon`, `:clearmemory`) - **Search integration** (`:search`, `:deepsearch`) Learn more in the [Inline CLI guide](https://apipie.ai/docs/features/inlinecli). --- ## Understanding State Management Controls In addition to core configuration settings, APIpie State Management includes **special control options** that determine how state can be changed, who it applies to, and how it persists. These controls make state even more powerful and adaptable across different use cases. ### Inline CLI Toggle The `enable_inline_cli` flag determines whether users can modify state directly from inside their prompts using the [Inline CLI](https://apipie.ai/docs/features/inlinecli). - **When enabled**, users gain the ability to naturally adjust models, memory, integrity, and other aspects of the API call inline with their query (`:setmodel:`, `:becreative`, etc.). - **When disabled**, all inline commands are ignored, and state can only be changed through the API or GUI. This toggle allows teams to decide how much flexibility to grant to end users. ### GUI-Only Mode The `gui_only` flag enforces that all state changes must be made in the APIpie dashboard. - **When enabled**, both API-based updates and Inline CLI commands are blocked. - **When disabled**, state can be updated via any method (API, CLI, or GUI). This is especially useful for enterprises that need strict governance and want to ensure changes are audited and centrally managed. ### Scope: App vs. User Every state configuration is tied to a **scope**, which can be either `app` or `user`. - **App-scoped state** applies to all requests made with a given API key. - **User-scoped state** applies individually to each `user` value sent with that API key. Each user maintains their own isolated state, making this ideal for **multi-tenant services** where every end-user experience needs to be personalized. Important notes about scope behavior: - An API key (and its associated app) can only support one scope: **either app-scoped or user-scoped, not both**. - Switching a state from `app` to `user` (or vice versa) will clear all existing configurations. - The `key` field in a state record reflects the `app_name` or API key name defined when the key was created. --- These special controls — Inline CLI, GUI-only mode, and scoping — ensure that state management works not only as a configuration system but also as a **flexible governance layer**. Together, they give individuals freedom when needed and enterprises the ability to enforce structure when required. --- ## Conclusion APIpie by Neuronic AI delivers **turnkey AI infrastructure** for developers and enterprises. With **Inline CLI for real-time prompt control**, **API for programmatic configuration**, and **GUI for enterprise governance**, you can manage state seamlessly across apps, API keys, and users. Start with Inline CLI (`:help`) or explore `/v1/state` for advanced control. # Preferred Routing: Optimize Cost or Performance ![Preferred Routing Feature Banner](https://apipie.ai/img/docs/features/routing-banner.svg){width="100%"} Discover the flexibility of our Preferred Routing feature, designed to give users control over how their AI requests are handled when multiple providers are available for a selected model. With this feature, users can specify routing based on price or performance, ensuring that their needs are always met, whether focused on cost efficiency or response speed. ## Why Use Preferred Routing? Preferred Routing offers a highly customizable approach to handling AI requests by allowing users to choose between the most cost-effective or the fastest provider. This is particularly useful for businesses that either want to minimize operational costs or prioritize performance for critical use cases. ## Understanding Preferred Routing Options ### Routing Types - **Price**: Automatically routes your request to the least expensive provider available for the selected model. Ideal for use cases where saving money is more important than response time. - **Perf**: (Default) Routes your request to the most responsive provider based on your prompt size, ensuring the fastest possible response. - **Perf-Avg**: Selects the provider with the best average latency across all prompt sizes, ensuring consistent response times in various scenarios. ## How to Use Preferred Routing Using Preferred Routing in your AI workflows is simple. Here's a typical API call with the routing parameter: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "messages": [ { "role": "system", "content": "why is the sky blue?" } ], "model": "gpt-3.5-turbo", "routing": "perf", //route to the most responsive provider of this model "temperature": 1, }' ``` This example demonstrates how to use the `routing` parameter in an API request. By default, the `routing` option is set to `perf`, but users can modify it to either `price` or `perf-avg` based on their needs. ## Benefits of Preferred Routing for Businesses By offering both price-based and performance-based routing, Preferred Routing allows businesses to customize their AI deployments to meet different operational goals: - **Cost Reduction**: Use price routing to minimize expenses for routine tasks where speed is not critical. - **Optimal Performance**: Leverage performance routing to ensure fast response times for time-sensitive tasks. ## Setting Up Preferred Routing To start using Preferred Routing effectively, follow these steps: 1. **Configure Your API Call:** - Add the `routing` parameter and set it to either `price`, `perf`, or `perf-avg` depending on your needs. 2. **Monitor Costs and Latency:** - Track your usage and adjust routing settings to balance cost savings and performance as needed. ## FAQs 1. **What is the default routing option?** - The default routing option is `perf`, which selects the most responsive provider based on your prompt size. 2. **How does price routing work?** - Price routing selects the cheapest provider for your request, based on real-time pricing for the model. 3. **Can routing impact response time?** - Yes, selecting price routing may result in slower response times, while performance routing will optimize for speed. 4. **Can I change the routing option after sending a request?** - No, the routing option must be set when making the API call. 5. **Is Preferred Routing available for all models?** - Preferred Routing is supported for models with multiple providers. If only one provider is available, routing options are not applicable. ## Links - [Chat Completions API](https://apipie.ai/docs/api/chatcompletions){rel=""nofollow""} - [Fetch Models API](https://apipie.ai/docs/api/fetchmodels){rel=""nofollow""} ## Conclusion Preferred Routing offers users the flexibility to prioritize either cost savings or performance when making AI requests. Whether your focus is on efficiency or speed, this feature ensures that your needs are met seamlessly across multiple providers. # Minimize AI Hallucinations with Integrity ![Integrity Feature Banner](https://apipie.ai/img/docs/features/integrity-banner.svg){width="100%"} Discover the power of our Integrity feature, specifically designed to enhance AI accuracy and reliability by minimizing hallucinations— a prevalent concern for businesses using AI technology. Integrity is ideal for applications where accuracy and dependability are paramount. ## Why Use Integrity? Integrity significantly reduces the likelihood of AI-generated errors, known as hallucinations, by implementing a robust validation process. This feature leverages the AI's ability to cross-verify answers to ensure you receive the most accurate response possible, making it invaluable for businesses requiring consistent and reliable AI responses. ## Understanding Integrity Settings ### Integrity Levels - **Integrity 11**: Default query with no additional checks. - **Integrity 12**: Queries the specified model twice, allowing AI to vote 5 times on the best answer. - **Integrity 13**: Queries the model thrice, selecting the best answer from three options. In both Integrity 12 and 13, the `gpt-4o` model is used by default for integrity checks, taking advantage of OpenAI's capability to provide `n` factor benefits for rapid cross-verification. ### How to Use Integrity Integrating Integrity into your AI projects is straightforward. Below is a typical API call focusing on the essential parameters for Integrity configuration: NOTE: Other than adding the integrity setting and the optional integrity model, there is no perceivable difference in the API call or its response, except for the response time. All the processes outlined in this document happen in the background. ## Example: Query Cohere Command on Openrouter using Open AI gpt-3.5 for Integrity ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "messages": [ { "role": "system", "content": "Why is the sky blue?" } ], "model": "command", "provider": "openrouter", "max_tokens": 300, "integrity": 12, "integrity_model": "gpt-3.5-turbo", }' ``` This code snippet demonstrates an API call utilizing Integrity to ensure the most reliable answer. Note that the `integrity_model` setting is optional and defaults to `gpt-4o`, but users may set it to any supported OpenAI model. The Response returned is an Open AI styled standard response. ## Benefits of Integrity for Businesses By mitigating potential hallucinations, the Integrity feature ensures that AI integrations remain dependable and accurate. This is particularly beneficial for business processes where data accuracy is crucial, such as report generation, customer query handling, and automated decision-making systems. **Note:** Integrity settings incur additional API costs due to repeated queries and voting mechanisms, but remain affordable and worthwhile, especially for high-stakes applications. ## Setting Up Integrity To effectively use Integrity in your AI workflows, follow these steps: 1. **Configure Your API Call:** - Set the `integrity` parameter to 12 or 13 depending on your accuracy needs. - Select the `integrity_model` only if deviating from the default `gpt-4o`. 2. **Understand the Cost and Time Implications:** - Ensure your application is suited to handle slight delays in response time due to the voting process. 3. **Test and Validate Results:** - Continuously monitor results to optimize the use of Integrity for your specific requirements. ## Tips & Tricks - **Leverage Integrity for Complex Queries:** Use Integrity for questions where accuracy is non-negotiable. - **Balance Cost and Accuracy:** Assess whether Integrity's additional cost aligns with your accuracy needs. - **Experiment with Different Models:** Although `gpt-4o` is the recommended starting point, trying other OpenAI models in many cases can yield similar results at a lower cost. ## FAQs 1. **What is the main purpose of the Integrity feature?** - To reduce hallucinations in AI responses, providing more reliable outputs. 2. **Why does Integrity cost more?** - The additional accuracy processes (i.e., multiple queries and voting) incur extra usage costs. 3. **Is Integrity suitable for all AI applications?** - It's best used for non-real-time applications where accuracy is prioritized over response speed. 4. **Can other AI models besides OpenAI be used for integrity checks?** - Currently, Integrity primarily supports OpenAI models due to specific feature support. 5. **How does Integrity affect response time?** - The response time is slightly longer due to added processing, so it is recommended for applications where this delay is acceptable. ## Links - [Chat Completions API](https://apipie.ai/docs/api/chatcompletions){rel=""nofollow""} - [Fetch Models API](https://apipie.ai/docs/api/fetchmodels){rel=""nofollow""} ## Conclusion With the completion of this guide, you now possess a comprehensive understanding of how to implement and benefit from the Integrity feature in your AI projects. We appreciate your commitment to leveraging Neuronic AI's tools and look forward to seeing how Integrity enhances your AI endeavors. # Inline CLI: Dynamic API Configuration in Prompts ![Inline CLI Feature Banner](https://apipie.ai/img/docs/features/inlinecli-banner.svg){width="100%"} Unlock powerful real-time API control with **Inline CLI** – a breakthrough feature built directly into your prompt flow. This innovation enables developers to dynamically apply API configurations within user input, empowering advanced AI behavior with a single line of text. ## What is Inline CLI? Inline CLI allows users to inject configuration and behavior commands into the **first or last 250 characters of any prompt** using the syntax `:command:value`. These commands are parsed before the prompt is sent to the model and can be used to persist state, override models, shape responses, and control memory, search, and more. Inline CLI works across any OpenAI-compatible API or agent. It’s powerful, portable, and frictionless for users of all levels. ## Authentication Use one of the following authentication methods: - Bearer token in the `Authorization` header - API key via `x-api-key` header Example headers: ```http Authorization: Bearer x-api-key: ``` ## Enabling Inline CLI Inline CLI is controlled via the `inline_cli` field in the request body. You can set: - `"all"` — Enable all CLI features - Comma-delimited features, e.g. `"model,search,memory"` - `"false"` — Fully disables Inline CLI Example: ```json "inline_cli": "model,shaping,integrity" ``` ## How Inline CLI Works - Commands prefixed with `:` (colon) are extracted from the prompt - They can be persistent (`set` commands) or temporary (for a single query) - Some commands (like `:help`, `:getstate`) bypass model processing - You can mix commands, like setting state and querying in one prompt --- ## Supported Commands ### 📄 Information Commands Direct responses (no prompt processing): | Command | Description | | ----------- | -------------------------------------- | | `:getstate` | Show current saved CLI settings | | `:help` | Show all available Inline CLI commands | --- ### 🤖 Model Commands Choose or override AI models: | Command | Description | | ------------------------------ | ---------------------------------------------------------- | | `:setmodel:/` | Persistently use specified model | | `:unsetmodel` | Remove saved model, revert to defaults | | `:answerwith` | Use a specific provider/model once (e.g. `:answerwithgpt`) | Supported AI aliases: `openai`, `gpt`, `claude`, `anthropic`, `grok`, `gemini`, `llama`, `deepseek`, `mistral`, `mixtral`, `smart`, `cheap` --- ### 🛠️ Shaping Commands Customize model behavior: | Command | Description | | ---------------- | -------------------------------- | | `:beprecise` | Lower temperature, more accurate | | `:bebalanced` | Balanced configuration | | `:becreative` | Higher creativity, more diverse | | `:becrazy` | Maximum randomness | | `:becoder` | Optimized for code responses | | `:avoidrepeat` | Penalize repeated tokens | | `:answerdiverse` | Increase answer diversity | | `:stayontopic` | Focus tightly on topic | --- ### ✅ Integrity Commands Eliminate hallucinations and ensure accurate responses: | Command | Description | | ----------------------- | --------------------------------- | | `:setintegrity` | Enable normal integrity setting | | `:setsuperintegrity` | Enable maximum integrity setting | | `:answerintegrity` | Use integrity override once | | `:answersuperintegrity` | Use super integrity override once | | `:unsetintegrity` | Remove persistent integrity | --- ### 🌐 Internet Search Commands Enrich prompts with real-time web search: | Command | Description | | ----------------------- | ---------------------------------------------- | | `:search` | Perform fast search | | `:searchmore` | Medium-depth search | | `:deepsearch` | Full-contextual search | | `:setsearchlang:` | Set search language (e.g. `:setsearchlang:en`) | | `:setsearchgeo:` | Set search region (e.g. `:setsearchgeo:US`) | --- ### 🧠 Memory Commands Persistent conversational memory: | Command | Description | | --------------------- | ----------------------------------------- | | `:setmemoryon` | Turn memory on | | `:setmemoryoff` | Turn memory off | | `:clearmemory` | Delete all memory for user/session | | `:setmemexpire:` | Set memory expiration in minutes (5-1440) | --- ## Inline CLI vs State Parameters | Feature | CLI `:command` | JSON Field | Behavior | | ------------------ | ---------------------------- | -------------------- | ---------------------- | | Persistent Setting | `:setmodel:gpt/4o` | `model`, `provider` | Stored in state | | One-Time Override | `:answerwithgpt` | N/A | Applies once | | View State | `:getstate` | N/A | No model call made | | Persist + View | `:setmodel:gpt/4o :getstate` | `inline_cli`, `user` | Set & view in one call | --- ## Example Prompt Usage ```bash curl -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Authorization: Bearer ' \ -H 'Content-Type: application/json' \ --data-raw '{ "user": "12345", "inline_cli": "all", "messages": [ { "role": "user", "content": "Tell me a fun fact :becreative :answerwithclaude" } ] }' ``` --- ## API Schema Integration Use the `ChatCompletionRequest` schema to configure: - `inline_cli`: `"all"` or comma-delimited feature list - `user`: Required for memory, CLI, and tenant-based features - `model`, `provider`: Can be overridden by inline CLI - `memory`, `mem_session`, `mem_clear`: Memory state and control - `rag_tune`, `search_geo`, `search_lang`: RAG and search customization - `tools`, `tool_choice`, `tools_model`: Optional function-calling support - `temperature`, `top_p`, `top_k`, `penalties`: Prompt shaping controls Full schema in API docs: :br[Chat Completions API](https://apipie.ai/docs/api/chatcompletions){rel=""nofollow""} --- ## Tips for Devs - You can mix `:set` and `:getstate` to both configure and review settings - If a user sends only info commands, the model isn’t invoked at all - All commands are ignored by the model and intercepted by the system --- ## Conclusion Inline CLI gives developers and users full control over AI behavior inside the natural language prompt. With built-in support across [memory](https://apipie.ai/docs/features/imm), [model control](https://apipie.ai/dashboard){rel=""nofollow""}, [search](https://apipie.ai/docs/features/internetsearch), [integrity](https://apipie.ai/docs/features/integrity), and more, it's a one-of-a-kind system designed to **maximize customization with zero overhead**. Start with `:help` in your prompt and build smarter agents with fewer constraints. --- ## Related Links ### Internal Documentation - [Chat Completions API](https://apipie.ai/docs/api/chatcompletions) - [Fetch Models API](https://apipie.ai/docs/api/fetchmodels) - [Ragtune API](https://apipie.ai/docs/features/ragtune) - [Vectors API](https://apipie.ai/docs/api/query-vectors) - [Internet Search Grounding](https://apipie.ai/docs/features/internetsearch) - [Model Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} - [Integrity](https://apipie.ai/docs/features/integrity) - [Integrated Model Memory](https://apipie.ai/docs/features/imm) # Global AI Operations Dashboard: Complete Model Analytics Platform ![Global AI Operations Dashboard](https://apipie.ai/img/docs/features/dashboard/dashboard-banner.png){width="100%"} Transform your AI model selection process with our comprehensive Global AI Operations Dashboard. Access real-time performance metrics, pricing analytics, and advanced filtering capabilities to find the perfect AI model for your needs. Whether you're seeking the **top AI API**, comparing **OpenAI competitors**, or need a **single AI** solution across multiple providers, our **API router dashboard** provides unparalleled insights into the **best AI APIs** available. ## Dashboard Overview The [APIpie Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} serves as your central command center for AI model analytics, providing comprehensive insights into hundreds of models across multiple providers. Our platform acts is a **one API** solution, giving you access to the **best AI models** from leading providers while offering detailed performance and cost analytics. **Browse models** easily with our intuitive interface to **list models** and discover **available models** across all providers. ### Why APIpie is the top router for AI APIs - **Real-world pricing analytics** - Not advertised rates, but where available actual **router costs** calculated from our proprietary algorithms - **Comprehensive latency tracking** across multiple token bucket sizes (2K, 4K, 8K, 16K, 32K tokens) - **Time to First Chunk (TTFC)** metrics for optimal **streaming APIs** performance - **Smart model grouping** for easy comparison across providers in our **model rankings** - **Advanced filtering and sorting** capabilities - **30-day availability tracking** with visual performance graphs for **better uptime** - **Multi-provider coverage** including OpenAI, Anthropic, Google, Meta, and open source alternatives ::note{to="https://apipie.ai/dashboard"} Visit our live Dashboard to explore all available AI models with real-time performance metrics and pricing analytics. :: ## Performance Metrics & Analytics Our dashboard provides the most comprehensive AI model analytics platform available, tracking performance across multiple dimensions to help you find the **top AI APIs** for your use case. Discover your **best LLM ranking** and **model ratings** to help you identify your **top models weekly** through our dashboard. ### Latency Tracking Across Token Buckets We measure and display latency performance across five different prompt/response size categories: - **2K tokens** - Short conversations and quick queries - **4K tokens** - Medium-length interactions and code generation - **8K tokens** - Complex reasoning and detailed analysis - **16K tokens** - Long-form content and document processing - **32K tokens** - Extended context and comprehensive analysis ### Time to First Chunk (TTFC) Metrics Critical for **streaming router** applications and real-time user experiences: - **Streaming responsiveness** - How quickly users see the first response from our **streaming APIs** - **Provider comparison** - TTFC performance across different providers - **Model optimization** - Identify the fastest models for interactive applications - **User experience optimization** - Choose models that provide the best perceived performance ### Real-World Pricing Analytics Unlike other platforms that show advertised rates, where available our pricing reflects actual **costs**: - **Proprietary cost algorithm** calculating real service costs per model/provider - **Per-million-token pricing** for transparent cost comparison - **Input and output pricing** separately tracked ## Advanced Filtering & Sorting Dashboard V2 introduces powerful filtering and sorting capabilities, making it the ultimate **ML & AI API** selection tool for developers and enterprises. **Browse models** efficiently with our filtering system. ### Filter by Model Types - **Language Models (LLMs)** - Text generation, reasoning, and conversation - **Voice Models** - Text-to-speech and audio processing - **Embedding Models** - Vector representations for search and similarity - **Image Generation** - Creative and artistic image creation ### Filter by Subtypes - **Multimodal** - Multimodal image analysis and Text OCR processing capabilities - **Chat** - Conversational AI optimized for dialogue - **Code** - Programming and software development focused - **Reasoning** - Advanced logical thinking and problem-solving - **Tools** - Function calling and external integration support - **& More** - See the live dashboard for further model subtypes ![AI Model Type Filters](https://apipie.ai/img/docs/features/dashboard/type-filters.png){width="50%"} ### Performance-Based Filtering - **Latency thresholds** - Exclude models above specific latency limits - **TTFC filtering** - Show only models meeting **streaming APIs** requirements - **Context window** - Filter by maximum token capacity ### Pricing Filters - **Maximum cost per million tokens** - Stay within budget constraints ![AI Model Filter Thresholds](https://apipie.ai/img/docs/features/dashboard/thresholds.png){width="50%"} ## Sorting Capabilities Organize models based on your priorities with our **model ranking** system: ### Performance Sorting - **Throughput** (ascending/descending) - Requests per minute capacity - **Latency** (ascending/descending) - Average response time - **TTFC** (ascending/descending) - **Streaming router** responsiveness ### Model Characteristics - **Newest** - Recently released models first - **Context size** - Maximum token capacity - **Model size** - Parameter count and model complexity ### Cost Analysis - **Input pricing** - Cost per million input tokens - **Output pricing** - Cost per million output tokens - **Image cost** - Standard Quality and size pricing ![AI Model Sorting](https://apipie.ai/img/docs/features/dashboard/sorting.png){width="35%"} ## AI Router API & Model Discovery Our **base API** and **router API** provide programmatic access to **browse models**, **list models**, and access **available models** across all providers. The **AI fetch** functionality enables seamless integration with your applications while our **streaming APIs** ensure optimal performance. ## Model Categories & Statistics Our dashboard provides comprehensive coverage of the AI ecosystem with detailed **AI stats**: ### Language Models - **Hundreds of models** across all major providers - **Text generation, reasoning, and conversation** - **Specialized coding models and technical models** - **Multilingual and domain-specific variants** ### Vision Models - **Multimodal models** with image understanding - **OCR and document analysis capabilities** - **Chart and diagram interpretation** - **Visual question answering and description** ### Voice Models - **Text-to-speech models** - **Multiple languages and voice styles** - **Real-time and batch processing options** - **High-quality audio generation** ### Specialized Models - **Image generation models** for creative applications - **Coding models** for software development with **coder** optimization - **Embedding models** for vector operations ![AI Overview](https://apipie.ai/img/docs/features/dashboard/overview.png){width="100%"} ## Provider Coverage Access the **best AI APIs** from leading providers through our **top ranked router** platform: ### Major Private Providers - **OpenAI** - GPT-4, GPT-3.5, and specialized models - **Google Gemini** - Advanced multimodal models and AI capabilities - **Anthropic** - Claude family and safety-focused models - **Perplexity** - Search-enhanced AI and reasoning models - **Amazon** - Nova models and cloud AI services - **& MORE** - Additional enterprise and specialized providers ### Open Source Alternatives - **Meta Llama models** - Open source language models - **Google Gemma** - Lightweight and efficient models - **Microsoft Phi** - Small but capable models - **Mistral models** - European AI excellence - **Community models & More** - Cutting-edge research implementations ## Real-Time Model Monitoring Stay informed with live performance tracking and **better uptime** monitoring: ### Availability Monitoring - **30-day availability statistics** for each model with **better uptime** tracking - **Real-time status indicators** showing current availability - **Historical uptime data** for reliability assessment - **Provider reliability comparison** across different services ### Performance Tracking - **Live latency updates** based on actual API calls - **TTFC monitoring** for **streaming APIs** performance - **Throughput capacity** tracking and limits - **Error rate monitoring** and quality metrics - **AI stats** for comprehensive performance analysis ### Cost Tracking - **Real-time pricing updates** reflecting actual - **Historical cost trends** and pricing changes - **Provider cost comparison** for identical models - **Budget impact analysis** for cost optimization ![AI Model Comparison](https://apipie.ai/img/docs/features/dashboard/grouped_models.png){width="100%"} ## Using the Dashboard Effectively ### Finding the Best AI Model 1. **Define your requirements** - Determine needed capabilities (text, vision, voice, etc.) 2. **Set performance criteria** - Establish latency and quality requirements 3. **Apply filters** - Narrow down options based on your needs 4. **Browse models** - Use our interface to **list models** and explore **available models** 5. **Compare grouped models** - Expand clusters to see provider comparisons 6. **Analyze metrics** - Review latency, cost, and availability data with **AI stats** 7. **Test desired models** - Use our **router API** to validate performance ### Optimizing for Cost-Performance - **Filter by budget constraints** to stay within spending limits - **Compare provider pricing** for identical model capabilities - **Consider token bucket performance** for your specific use case ### Ensuring Reliability - **Check 30-day availability** statistics for each model with **better uptime** metrics - **Review provider reliability** across different services - **Monitor real-time status** before critical deployments - **Track API limitations** to avoid service interruptions ## API Access to Dashboard Data Access dashboard insights programmatically through our **base API** and **router API**: - **AI fetch** capabilities for real-time model data - **Streaming APIs** for live performance metrics - **AI stats** and analytics data access ### Cost Management - **Monitor pricing trends** for budget planning - **Compare provider costs** for identical capabilities - **Consider usage patterns** when selecting models - **Use cost-performance ratios** for value optimization --- ## Getting Started 1. **Visit the Dashboard** at [apipie.ai/dashboard](https://apipie.ai/dashboard){rel=""nofollow""} 2. **APIpie Router login** - Create an account for advanced features 3. **Explore model categories** to understand **available models** 4. **Browse models** using our filtering system 5. **Apply filters** based on your requirements 6. **Compare models** within groups for best options 7. **Review performance metrics** with **AI stats** for your selected models 8. **Start testing** with our **router API** integration The APIpie Dashboard represents the most comprehensive **AI API** analytics platform available, serving as your **top router** for AI model selection. Our **streaming router** provides the insights you need to make informed decisions about **AI and ML** model selection. Whether you're comparing **OpenAI competitors**, seeking **open source AI API** alternatives, or need a reliable **top ranked router** platform, our dashboard provides the data-driven insights essential for optimal AI implementation. Experience the future of AI model selection and management with the APIpie Global AI Operations Dashboard - your **top router** for AI APIs. # Internet Search Grounding: Web Data in AI Responses ![Search Feature Banner](https://apipie.ai/img/docs/features/search-banner.svg){width="100%"} Enable any model to answer with [real-time, verifiable information](https://arxiv.org/abs/2302.13007){rel=""nofollow""} using **Inline Search** — our built-in web search and grounding system that works seamlessly with OpenAI-compatible APIs. With a single request, you can inject live context from the internet into your prompt, no browser tools, agents, or plugins required. ## What is Search Grounding? Inline Search enhances any model by integrating real-time search results and scraped content directly into the prompt — without requiring the model to have built-in browsing capabilities. When enabled, the system: 1. Performs X number of live searches using top-ranked search engines 2. Fetches and ranks up to 100 results per search 3. Scrapes the top results, cleans the text, removes noise 4. Appends that content to your prompt before model invocation This allows even legacy models to answer with fresh knowledge, and it works **automatically with any model or app built on the OpenAI chat format.** --- ## Why Use Inline Search? - **Fresh data in any model** – including GPT-3.5, Claude, LLaMA, Mixtral, and more - **No special setup** – works out of the box in standard OpenAI API format - **Customize search volume and scope** - **Simple to configure via request body** > ⚠️ *Note*: Because this uses live search + scraping, expect an added delay of 1–6 seconds total, depending on complexity of the web pages we are pulling data from. We have to render the javascript before we can pull the data. --- ## Quick Start (OpenAI Format Example) Use `web_search_options` in your request to instantly enable Inline Search. ```bash curl -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Authorization: Bearer ' \ -H 'Content-Type: application/json' \ -d '{ "user": "qa_test", "stream": true, "model": "openai/gpt-4o", "web_search_options": { "search_context_size": "medium" }, "max_tokens": 500, "messages": [ { "role": "user", "content": "Who was the best and worst president of the united states, quantifiable and verifiable?" } ] }' ``` ### Options for `web_search_options.search_context_size` | Value | Description | | -------- | --------------------------------------------------------- | | `low` | Small content injection (\~1 result up to 10k characters) | | `medium` | Moderate (\~3 results, \~15K characters) | | `high` | Full web grounding (\~5 results, \~35K characters) | --- ## Alternate Inline Search via Payload (Advanced) You can also enable search with deeper control using the `online` flag and extended parameters: | Field | Description | | --------------- | -------------------------------------------------------------- | | `online` | Enable inline search grounding (`true` or `false`) | | `searches` | Number of search queries to perform (default: 1) | | `pull` | Max results to pull from each search (default: 20) | | `use` | How many of those results to append to the prompt (default: 3) | | `scrape_length` | Max number of characters to include from each scrape | | `search_lang` | Language code for results (e.g. `"en"`) | | `search_geo` | Geolocation code (e.g. `"US"`) | ```json { "user": "user123", "model": "claude-3-5-sonnet", "online": true, "searches": 2, "pull": 10, "use": 4, "scrape_length": 12000, "search_lang": "en", "search_geo": "US", "messages": [{ "role": "user", "content": "Summarize the most recent news about generative AI in healthcare." }] } ``` --- ## Using Inline CLI (Prompt Commands) You can also enable search directly in the prompt with these Inline CLI commands: | Command | Behavior | | ------------------- | ------------------------------------------ | | `:search` | Fast search (1 query, 1 results) | | `:searchmore` | More results (2 query, 3 results) | | `:deepsearch` | Broad + deep search (3 queries, 5 results) | | `:setsearchlang:en` | Set language to English | | `:setsearchgeo:US` | Set geo location to United States | Example prompt: ```json { "content": "Summarize today’s major AI headlines :deepsearch :setsearchlang:en :setsearchgeo:US" } ``` --- ## Direct Use of the Search & Scrape APIs For users needing lower-level control or standalone search capabilities, we also offer public API endpoints. ### POST `/v1/search` ```json { "query": "latest AI developments", "search_provider": "google", "search": "google", "geo": "us", "lang": "en", "results": 10, "safeSearch": -1, "user": "user123" } ``` Returns a list of ranked search results with URLs, titles, and descriptions. --- ### POST `/v1/scrape` ```json { "url": "https://example.com/article", "format": "parsed" } ``` Returns parsed content: - `title` - `textContent` - `excerpt` - `hrefs[]` --- ## Developer Tips - Works with **all models**, including offline-only models like GPT-3.5 or Mistral - Use `web_search_options` for OpenAI-style apps - Use `online`, `pull`, `use`, `scrape_length` for full control - Combine with `memory`, `integrity`, and `shaping` for robust agents - Monitor latency: search grounding can add 1–6s per request --- ## Example: High-Precision Search ```bash curl -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Authorization: Bearer ' \ -H 'Content-Type: application/json' \ --data-raw '{ "user": "research_bot", "model": "mixtral", "online": true, "searches": 3, "pull": 15, "use": 5, "scrape_length": 15000, "search_lang": "en", "search_geo": "US", "messages": [ { "role": "user", "content": "What are the leading companies developing open-source LLMs in 2025?" } ] }' ``` --- ## Conclusion Inline Search brings live, trusted information to **any model or application** using the standard OpenAI request format or advanced options. Whether you're building intelligent agents, research assistants, or chatbots — this feature ensures your responses are grounded in the real world. Try it out with `web_search_options`, `online: true`, or [`:deepsearch`](https://apipie.ai/docs/features/inlinecli) in your next request. --- ## Related Links ### Internal Documentation - [Chat Completions API](https://apipie.ai/docs/api/chatcompletions) - [Web Search API](https://apipie.ai/docs/api/web-search) - [Web Scrape API](https://apipie.ai/docs/api/web-scrape) - [Inline CLI Commands](https://apipie.ai/docs/features/inlinecli) - [RAG Integration](https://apipie.ai/docs/features/ragtune) - [Vector Database Integration](https://apipie.ai/docs/api/query-vectors) - [Integrity Checks](https://apipie.ai/docs/features/integrity) - [Model Selection](https://apipie.ai/docs/features/models) # AI Model Pooling: Enhance Reliability & Security ![Pooling Feature Banner](https://apipie.ai/img/docs/features/pooling-banner.svg){width="100%"} Unleash the potential of AI Model Pooling, a breakthrough feature designed to enhance the reliability, redundancy, and data security of AI-generated outputs. This guide delves into the essence of Pools, a service that aggregates multiple similar models to ensure optimal performance and service continuity, especially beneficial for large-scale and sensitive AI applications. ## Why Use AI Model Pooling? AI Model Pooling offers higher rate limits, reliable redundancy, and improved data security. By spreading requests across multiple providers and similar models in a pool, it ensures consistent response delivery. This method not only enhances reliability but also minimizes the exposure of sensitive data to any single AI provider. AI Model Pooling is particularly useful for applications where high availability and data privacy are crucial. ## Understanding AI Model Pooling Settings ### Pool Structures - **Model Aggregation:** Each pool contains several similar models that provide higher rate limits and ensure redundancy. - **Automatic Failover:** If a request fails due to a provider error, the pool will automatically reroute the query to another member until a successful response is obtained. - **Context Length Value:** The suffix value (e.g., 4k, 8k) indicates the maximum response tokens supported by the models within a pool. Below is an API example to demonstrate how to leverage model pooling: ```bash curl -L -X GET 'https://apipie.ai/v1/models?subtype=pool' \ -H 'Authorization: Bearer ' \ -H 'Content-Type: application/json' \ { "object": "list", "data": [ { "type": "llm", "subtype": "pool", "provider": "pool", "model": "gpt-3.5_4k", "description": "A pool of like models for higher rate limits, reliable redundancy and improved data security", "max_response_tokens": 4096 } ] } ``` Now we can use any model from that list in a regular chat completions format like this > ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "messages": [ { "role": "system", "content": "Why is the sky blue?" }, "provider": "pool", "model": "gpt4o_16k" }' ``` This API call showcases how to list the available pools, enabling optimized selection based on your application requirements. ## Benefits of AI Model Pooling for Businesses Leverage AI Model Pooling for enhanced reliability and data privacy across AI applications. By ensuring a seamless and dependable service, Pools make it possible to maintain service standards even during peak times or unforeseen provider issues. ## How to Set Up AI Model Pooling To effectively integrate Pools into your AI workflows, follow these steps: 1. **Configure Your API Call:** - Change the `provider` to "pool". - Use the desired pool name in the `model` field. 2. **Access the Model List:** - Use the models route with filter `subtype=pool` to retrieve available pools and their details. 3. **Test and Validate Results:** - Continually monitor outcomes to ensure the pooling mechanism meets your application’s needs. ## Tips & Tricks - **Optimize Requests:** Use Pools to handle higher request volumes or when you need a response every time, without fail. - **Greater Privacy:** Distribute data exposure across multiple models and providers. - **Select Appropriate Pools:** Match pools based on model capabilities and response token limits to your use case needs. ## FAQs 1. **How does Model Pooling enhance reliability?** - By rerouting failed requests across multiple similar models on failure, ensuring a consistent response. 2. **Are there any additional costs associated with Pooling?** - No, there is no additional cost. Various providers may charge different rates for the same model, but there are no hidden additional fees beyond standard usage rates. 3. **Can I use Model Pooling with any AI application?** - While most applications can benefit, it is especially useful where reliability and data security are priorities. 4. **Is there any impact on response time when using Pools?** - Failover responses are swift, typically with negligible impact on perceived user latency. 5. **How to verify available Pools?** - Utilize the models API endpoint with relevant filters to view current pools. 6. **Do pools support streaming?** - Yes, all pools support streaming. Note that some Bedrock models may aggregate before streaming. Use the ChatX pools for guaranteed full streaming and memory support. ## Links - [Chat Completions API](https://apipie.ai/docs/api/chatcompletions){rel=""nofollow""} - [Fetch Models API](https://apipie.ai/docs/api/fetchmodels){rel=""nofollow""} - [Ragtune API](https://apipie.ai/docs/api/ragtune){rel=""nofollow""} - [Vectors API](https://apipie.ai/docs/api/vectors){rel=""nofollow""} ## Conclusion By using AI Model Pooling, you can ensure optimal service reliability and data security in your AI projects. We encourage you to implement Pooling to maximize the benefits and efficiencies of your AI applications. Thank you for choosing Neuronic AI, your trusted AI solutions partner. # Multimodal Vision Guide: AI Image Analysis & Text Extraction ![Multimodal Vision Feature Banner](https://apipie.ai/img/docs/features/vision-banner.svg){width="100%"} Transform your applications with advanced visual AI capabilities. Our Multimodal Vision API enables you to analyze images, extract text, understand visual content, and generate detailed descriptions. Whether you're building OCR systems, content moderation tools, or accessibility features, our vision API provides powerful image understanding across multiple AI providers. ## Multimodal Vision Overview The Multimodal Vision API allows developers to send images alongside text prompts to AI models that support visual understanding. These models can analyze images, extract text (OCR), describe visual content, answer questions about images, and perform complex visual reasoning tasks. ### OpenAI Compatible Framework Our Vision API maintains full compatibility with OpenAI's vision format while extending support to multiple providers. Simply include images in your chat completion requests using either image URLs or base64-encoded data, and the model will process both text and visual information together. ::note{to="https://apipie.ai/dashboard"} Check our [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} to see the complete list of multimodal-capable models and their vision features. :: ## How Multimodal Vision Works Vision-enabled models process both text prompts and images simultaneously, allowing for sophisticated visual understanding tasks. You can include images in your conversation by adding them to message content using either: - **Image URLs**: Direct links to publicly accessible images - **Base64 encoding**: Embedded image data within the request ::note - Images can be in various formats: JPEG, PNG, GIF, WebP - Maximum image size varies by provider (typically 4-20MB) - Some models support multiple images in a single request - All vision capabilities use the same chat completions endpoint :: ### Available Vision-Enabled Models Our API supports vision capabilities across multiple providers: - [OpenAI](https://apipie.ai/docs/models/openai) - GPT-4o, GPT-4 Vision - [Anthropic](https://apipie.ai/docs/models/claude) - Claude 3 series with vision - [Google](https://apipie.ai/docs/models/google) - Gemini Pro Vision, Gemini Ultra - [Meta](https://apipie.ai/docs/models/llama) - Llama Vision models Visit our [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} to explore all multimodal-capable models and their specific vision features. ## Choosing the Right Vision Model Selecting the optimal vision model depends on your specific use case, performance requirements, and budget considerations: ### Performance vs Cost Tiers - **Premium Models**: GPT-4o, Claude 3.5 Sonnet, Gemini Pro Vision - Best for complex visual reasoning and detailed analysis - Higher accuracy but increased cost - Ideal for professional applications requiring high precision - **Balanced Models**: GPT-4o-mini, Gemini Flash, Claude Haiku - Good performance at moderate cost - Suitable for most production applications - Excellent for general-purpose vision tasks - **Budget-Friendly Models**: Smaller vision models, open-source alternatives - Cost-effective for high-volume processing - Basic vision capabilities - Good for simple OCR and image description tasks ### Model Strengths by Use Case - **Text Extraction (OCR)**: GPT-4o, Claude Sonnet models excel at extracting and formatting text from complex documents - **Detailed Image Analysis**: Gemini Pro Vision and GPT-4o provide comprehensive scene understanding - **Technical Diagrams**: Claude models perform well with charts, graphs, and technical drawings - **Multiple Images**: Some models support comparing multiple images in a single request - **Speed-Critical Applications**: Gemini Flash and GPT-4o-mini offer faster response times ### Evaluation Tips 1. **Test with your data**: Use the [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} to test different models with your specific image types 2. **Consider context length**: Some models handle longer conversations with images better 3. **Check language support**: Ensure the model supports your required languages for OCR tasks 4. **Monitor costs**: Use smaller models for development and scale up for production 5. **Leverage routing**: Use our intelligent routing to automatically select optimal models ### Example Vision API Call with Image URL ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "model": "gpt-4o", "max_tokens": 300, "messages": [ { "role": "user", "content": [ { "type": "text", "text": "What do you see in this image? Describe it in detail." }, { "type": "image_url", "image_url": { "url": "https://example.com/sample-image.jpg" } } ] } ] }' ``` ### Example Vision API Call with Base64 Image ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "model": "gpt-4o", "max_tokens": 300, "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Extract all text from this document image." }, { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABAAD/2wBDAAYEBQYFBAYGBQYHBwYIChAKCgkJChQODwwQFxQYGBcUFhYaHSUfGhsjHBYWICwgIyYnKSopGR8tMC0oMCUoKSj/2wBDAQcHBwoIChMKChMoGhYaKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCj/wAARCAABAAEDASIAAhEBAxEB/8QAFQABAQAAAAAAAAAAAAAAAAAAAAv/xAAUEAEAAAAAAAAAAAAAAAAAAAAA/8QAFQEBAQAAAAAAAAAAAAAAAAAAAAX/xAAUEQEAAAAAAAAAAAAAAAAAAAAA/9oADAMBAAIRAxEAPwCdABmX/9k=" } } ] } ] }' ``` ### Response Example The response structure is identical to regular chat completions but includes visual analysis: ```json { "id": "chatcmpl-vision-5fde5f7fffe8d6dc1f18aab4a138d4b7", "object": "chat.completion", "created": 1729535643, "provider": "openai", "model": "gpt-4o", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "I can see a document containing several paragraphs of text. The document appears to be a business report with the following visible text:\n\n'QUARTERLY SALES REPORT\nQ3 2024 Performance Summary\n\nSales increased by 15% compared to Q2 2024...\n\nThe document includes charts showing monthly trends and appears to be professionally formatted with headers and structured content." }, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 1150, "completion_tokens": 85, "total_tokens": 1235, "prompt_characters": 45, "response_characters": 312, "cost": 0.01435, "latency_ms": 3420 } } ``` ## Vision-Specific Parameters ### Image Content Structure When including images in your messages, use this content structure: ```json { "role": "user", "content": [ { "type": "text", "text": "Your text prompt here" }, { "type": "image_url", "image_url": { "url": "https://example.com/image.jpg" } } ] } ``` ### Multiple Images Some models support analyzing multiple images in a single request: ```json { "role": "user", "content": [ { "type": "text", "text": "Compare these two images and describe the differences." }, { "type": "image_url", "image_url": { "url": "https://example.com/image1.jpg" } }, { "type": "image_url", "image_url": { "url": "https://example.com/image2.jpg" } } ] } ``` ## Common Vision Use Cases ### Text Extraction (OCR) Extract text from documents, signs, screenshots, or any image containing text: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "model": "gpt-4o", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Extract all text from this image and format it as clean, readable text." }, { "type": "image_url", "image_url": { "url": "https://example.com/document.jpg", "detail": "high" } } ] } ] }' ``` ### Image Description and Analysis Generate detailed descriptions of images for accessibility or content understanding: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "model": "claude-3-5-sonnet", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Provide a detailed description of this image for visually impaired users. Include colors, objects, people, activities, and spatial relationships." }, { "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } } ] } ] }' ``` ### Visual Question Answering Ask specific questions about image content: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "model": "gemini-pro-vision", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "How many people are in this image? What are they wearing? What is the setting?" }, { "type": "image_url", "image_url": { "url": "https://example.com/group-photo.jpg" } } ] } ] }' ``` ### Document Analysis Analyze charts, graphs, tables, and structured documents: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "model": "gpt-4o", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Analyze this chart and provide a summary of the key trends and data points." }, { "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mP8/5+hHgAHggJ/PchI7wAAAABJRU5ErkJggg==" } } ] } ] }' ``` ## Image Format Support ### Supported Formats (across most models) - **JPEG**: Most common format, good compression - **PNG**: Supports transparency, lossless compression - **GIF**: Animated images (first frame analyzed) - **WebP**: Modern format with excellent compression ### Size Limitations - Maximum file size varies by provider (4MB - 20MB) - Recommended resolution: 2048x2048 pixels or smaller - Higher resolution images may be automatically resized ### Base64 Encoding For base64 images, use the data URL format: ```text data:image/jpeg;base64, ``` Example Python code to encode an image: ```python import base64 def encode_image(image_path): with open(image_path, "rb") as image_file: return base64.b64encode(image_file.read()).decode('utf-8') # Usage base64_image = encode_image("path/to/your/image.jpg") data_url = f"data:image/jpeg;base64,{base64_image}" ``` ## Usage Metrics and Costs Vision requests typically use more tokens than text-only requests due to image processing: ### Token Usage - Images are converted to tokens for processing - Token count depends on image size ### Cost Optimization - Resize images to optimal dimensions before uploading - Consider image compression to reduce file size - Use caching for repeated image analysis Example response with vision usage metrics: ```json { "usage": { "prompt_tokens": 1150, "completion_tokens": 85, "total_tokens": 1235, "prompt_characters": 45, "response_characters": 312, "cost": 0.01435, "latency_ms": 3420 } } ``` ## Best Practices ### Image Quality - Use high-resolution images for better text recognition - Ensure good lighting and contrast in photos - Avoid blurry or distorted images - For documents, scan rather than photograph when possible ### Prompt Engineering - Be specific about what you want to extract or analyze - Use clear, descriptive prompts - Ask for structured output when needed (JSON, tables, lists) - Provide context about the image type (document, photo, chart, etc.) ### Error Handling Common vision-specific errors: - **Unsupported image format**: Check image format and convert if needed - **Image too large**: Reduce image size or compression - **Invalid image url**: Verify URL accessibility and format - **Model doesnt support vision**: Verify URL accessibility and format ### Security Considerations - Validate image URLs before processing - Sanitize extracted text for security - Be aware of privacy implications when processing images - Use HTTPS URLs for image references - Consider data retention policies for processed images --- ## Getting Started 1. **Choose a vision-enabled model** from our [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} 2. **Prepare your images** in supported formats (JPEG, PNG, GIF, WebP) 3. **Structure your API request** with both text and image content 4. **Test with simple use cases** like basic image description 5. **Optimize for your specific needs** using appropriate prompts Our Multimodal Vision API opens up powerful possibilities for AI image to text conversion, visual understanding, and document analysis. Start building your vision-powered applications today! # Tools Support: Optimize Your AI Integration Today ![Tools Support Feature Banner](https://apipie.ai/img/docs/features/tools-banner.svg){width="100%"} Explore the flexibility and power of our Tools Support feature, enabling seamless integration with both OpenAI and Anthropic models to handle a wide array of tool-based queries. This feature enhances AI capabilities, giving businesses the freedom to choose how tool requests are processed. ## Why Use Tools Support? Tools Support allows you to configure which model responds to your tool queries. You can select any model from either OpenAI or Anthropic to handle tool-based requests, ensuring you get the most effective response depending on your specific use case. The feature supports all models available on our platform and provides the flexibility to customize how your tool queries are handled. ## Understanding Tools Configuration ### ToolsModel Selection - **Default Model**: If no specific tools model is selected, the default `gpt-4.1-nano` is used. This model is optimized for efficiency and accuracy, offering reliable tool responses at a lower cost. - **Custom Model**: You can select any OpenAI or Anthropic model to handle tool calls. This is useful when you want to leverage specific model strengths for particular tools, such as weather updates or information retrieval. When using a different model for tools, we first send your request to the tools model with minimal `max_tokens`. If the response includes a tool call, we process it and return the tool's response. If no tool is called, we send the request to your chosen model to continue processing. ### Tools Query Process If your tool and query models are the same, the entire process happens within a single query. If the tool model is different, your request is divided into two phases: - **Phase 1**: A query is sent to the tools model to check for any tool invocations. - **Phase 2**: If the tool model doesn't return a tool response, the request is forwarded to your chosen model for standard processing. In cases where tools are invoked, you only pay for the tool query cost and the final model response. We optimize token usage by minimizing token consumption for tool checks. ### Example API Call Using Tools Here’s how you can integrate tools in an API call: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "messages": [ { "role": "user", "content": "what is the weather in dallas and what are some fun facts about Dallas?" } ], "model": "gpt-4o", "provider": "openrouter", "tools": [ { "type": "function", "function": { "name": "get_current_weather", "description": "Get the current weather in a given location", "parameters": { "type": "object", "properties": { "location": { "type": "string", "description": "City and state for the weather report, e.g., Dallas, TX" }, "unit": { "type": "string", "enum": ["celsius", "fahrenheit"] } }, "required": ["location", "unit"] } } }, { "type": "function", "function": { "name": "fun_city_facts", "description": "Get interesting facts about a city", "parameters": { "type": "object", "properties": { "location": { "type": "string", "description": "City to learn fun facts about, e.g., Dallas, TX" } }, "required": ["location"] } } } ], "tool_choice": "auto", "tools_model": "gpt-4o-mini" }' ``` ### Example Response: ```json { "id": "chatcmpl-8de1a7f199e44f616238f6f6dbf7dcf9", "object": "chat.completion", "created": 1729442924, "model": "gpt-4o", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Please bear with me for a moment while I gather the current weather information and some fun facts about Dallas for you.", "tool_calls": [ { "id": "call_vEOdJs1QG6zqvzVjYqA1Dl3h", "type": "function", "function": { "name": "get_current_weather", "arguments": "{\"location\": \"Dallas, TX\", \"unit\": \"fahrenheit\"}" } }, { "id": "call_gGlaXomXwVpThaH2ol6CqmfB", "type": "function", "function": { "name": "fun_city_facts", "arguments": "{\"location\": \"Dallas, TX\"}" } } ] }, "logprobs": null, "finish_reason": "tool_calls" } ], "usage": { "prompt_tokens": 159, "completion_tokens": 80, "total_tokens": 239, "prompt_characters": 194, "response_characters": 310, "cost": "0.00199500", "latency_ms": 1791 }, "system_fingerprint": "fp_d1ab9353" } ``` This API request demonstrates how to specify tool invocations in your query, and the response shows how tools are invoked and handled by the system. ## Benefits of Tools Support for Businesses The Tools Support feature provides unmatched flexibility and accuracy for handling tool-based queries. By allowing different models for tools, businesses can optimize costs and performance for specific use cases, such as: - Real-time data retrieval, like weather reports or fact-checking. - Dynamic integration of functions for improved user experiences. - Reduced overhead by utilizing the most cost-effective models for non-critical queries. **Note:** If the model designated for tools does not return a tool response, additional API charges apply for forwarding the query to the user's primary model. ## Setting Up Tools Support To effectively use Tools Support in your AI workflows, follow these steps: 1. **Choose Your Tools Model:** - Set the `tools_model` parameter to any supported OpenAI or Anthropic model based on your needs. - Use the default `gpt-4o-mini` for cost-effective and reliable responses. 2. **Configure Tool Query Settings:** - Ensure your API request includes the appropriate tool-related parameters and functions. 3. **Monitor and Optimize:** - Track the performance of tool queries to assess their impact on response times and costs. ## Tips & Tricks - **Leverage Tools for Dynamic Data**: Use tools for real-time data retrieval, like weather and factual data. - **Customize Your Tools Model**: Adjust the tools model to match the complexity of your queries. - **Optimize Token Usage**: Keep `max_tokens` low for tool queries to reduce costs. ## FAQs 1. **What is the default tools model?** - The default tools model is `gpt-4o-mini`, which balances accuracy and cost-effectiveness. 2. **Can I use Anthropic models for tools queries?** - Yes, both OpenAI and Anthropic models are supported for tool queries. 3. **What happens if the tools model doesn't return a tool response?** - The query is sent to your primary model for standard processing, incurring an additional charge. 4. **Is there any delay when using tools?** - There may be a slight delay when streaming tool responses due to the additional data processing. ## Links - [Chat Completions API](https://apipie.ai/docs/api/chatcompletions){rel=""nofollow""} - [Fetch Models API](https://apipie.ai/docs/api/fetchmodels){rel=""nofollow""} ## Conclusion With Tools Support, you're equipped to make the most of AI's ability to integrate external functions seamlessly. By customizing the model that handles tool queries, you can enhance your AI's capabilities while maintaining control over costs and performance. # Integrated Model Memory (IMM): Conversation Mgmt ![Integrated Model Memory (IMM) - Advanced AI Memory Management System](https://apipie.ai/img/docs/features/IMM.svg){width="100%"} # Integrated Model Memory (IMM) ## Overview Integrated Model Memory (IMM) is our implementation of [Cache Augmented Generation (CAG)](https://github.com/hhhuang/CAG){rel=""nofollow""}, designed to revolutionize how AI applications handle conversation context. Unlike traditional CAG systems that are typically tied to specific models, IMM provides a model-independent memory solution that works seamlessly across [all supported AI models](https://apipie.ai/docs/Models/Overview.md). This means your conversation history and context persist within a session regardless of which model you use - you can start a conversation with [GPT-4](https://platform.openai.com/docs/models/gpt-4){rel=""nofollow""}, continue with [Claude](https://docs.anthropic.com/claude/docs){rel=""nofollow""}, and switch to [Mistral](https://docs.mistral.ai/){rel=""nofollow""} while maintaining complete context. This unique approach eliminates the complexity of manual memory management while providing intelligent context retention across conversations. As our enhanced implementation of CAG, IMM simplifies AI development by automating critical memory-related tasks: - Automatic context management - Seamless conversation history tracking - Intelligent memory expiration handling - Multi-session support for different users or use cases ## Key Features and Benefits ### Effortless Implementation - Enable memory with a simple `"memory": 1` parameter - No need for complex vector database management - Automatic context retention and retrieval ### Advanced Session Management - Create independent memory instances with `mem_session` - Perfect for multi-user applications - Isolated conversation contexts for different use cases ### Smart Memory Controls - Configure memory expiration with `mem_expire` - Automatic cleanup of outdated conversations - Efficient resource management ### Developer-Friendly Design - Simple API integration - Flexible configuration options - Model-independent memory management - Persistent session memory across different models - Intelligent caching and retrieval mechanisms --- ## Implementation Guide ### Enabling Memory Management Activate IMM by including the memory parameter in your API request: ```bash {6-8} curl https://apipie.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer " \ -data-raw '{ "stream": true, "memory": 1, "mem_expire": 120, "mem_session": "test123", "provider": "openai", "model": "gpt-4o", "max_tokens": 100, "messages": [ { "role": "user", "content": "where did I put my car keys? I already forgot." } ] }' ``` ### Verifying Memory Functionality To see Integrated Model Memory in action, follow this sequence: #### Step 1: Save Information to Memory ```bash {5-7} curl -X POST "https://apipie.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer " \ --data-raw '{ "memory": 1, "mem_expire": 60, "mem_session": "test123", "provider": "openai", "model": "gpt-4o", "max_tokens": 100, "messages": [ { "role": "user", "content": "I put my car keys in the coffee can by the swing on the front porch." }, { "role": "assistant", "content": "Okay, I will keep your secret safe." }, { "role": "user", "content": "Why is the sky blue?" }, { "role": "assistant", "content": "The sky appears blue due to a phenomenon called Rayleigh scattering..." }, { "role": "user", "content": "What did I just ask you?" } ] }' ``` #### Step 2: Ask a Follow-up Question Now, ask the AI where you put your car keys and note how you do not need to include the message history: ```bash {5-6} curl -X POST "https://apipie.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer " \ --data-raw '{ "memory": 1, "mem_session": "cross_model_test", "provider": "openai", "model": "gpt-4o", "max_tokens": 50, "messages": [ { "role": "user", "content": "Where did I put my car keys?" } ] }' ``` #### Expected Response ```json { "id": "chatcmpl-6008a7442651dc5b433b58db12fa88ee", "object": "chat.completion", "created": 1737469664, "provider": "openai", "model": "gpt-4o", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "You put your car keys in the coffee can by the swing on the front porch." }, "logprobs": null, "finish_reason": "stop" } ] } ``` This confirms that the AI correctly remembers past conversations and can recall stored information. ### Cross-Model Memory Example One of IMM's unique features is its ability to maintain context across different AI models within the same session. Here's how to switch models while preserving conversation context: # Start with GPT-4 ```bash {7-8} curl -X POST "https://apipie.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer " \ --data-raw '{ "memory": 1, "mem_session": "cross_model_test", "provider": "openai", "model": "gpt-4o", "messages": [ { "role": "user", "content": "My favorite color is blue." } ] }' ``` # Continue with Claude, maintaining context ```bash {7-8} curl -X POST "https://apipie.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer " \ --data-raw '{ "memory": 1, "mem_session": "cross_model_test", "provider": "anthropic", "model": "claude-2", "messages": [ { "role": "user", "content": "What's my favorite color?" } ] }' ``` # Switch to Mistral, still maintaining context ```bash {7-8} curl -X POST "https://apipie.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer " \ --data-raw '{ "memory": 1, "mem_session": "cross_model_test", "provider": "mistral", "model": "mistral-large", "messages": [ { "role": "user", "content": "Can you confirm my favorite color?" } ] }' ``` Each model will have access to the full conversation history, demonstrating IMM's unique ability to maintain context across different AI models. --- ### Managing Multiple Sessions Create separate memory contexts for different users or applications, separate the sessions by differentiating the mem\_session string. ```bash {7} curl -X POST "https://apipie.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer " \ --data-raw '{ "memory": 1, "mem_expire": 60, "mem_session": "user456", "provider": "openai", "model": "gpt-4o", "max_tokens": 100, "messages": [ { "role": "user", "content": "donde esta la bibliteca?" } ] }' ``` ### Memory Expiration Control Set custom expiration times for conversation contexts to a maximum 1440 min with a default of 15: ```bash {6} curl -X POST "https://apipie.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer " \ --data-raw '{ "memory": 1, "mem_expire": 120, "mem_session": "timed_session", "provider": "openai", "model": "gpt-4o", "messages": [ { "role": "user", "content": "This conversation will expire in 120 minutes." } ] }' ``` ### Clearing Memory Clear specific session memory when needed, If you do not specify a session the default unnammed session will be cleared: ```bash {6} curl -X POST "https://apipie.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer " \ --data-raw '{ "memory": 1, "mem_clear": 1, "mem_session": "session_to_clear", "provider": "openai", "model": "gpt-4o", "messages": [ { "role": "user", "content": "Clear this session's memory." } ] }' ``` Expected response: ```json { "queryResults": "Memory session deleted successfully", "queryDetails": { "provider": "openai", "route": "gpt-4o", "promptTokens": 0, "responseTokens": 0, "promptChar": 0, "responseChar": 31, "cost": 0, "latencyMs": 0 } } ``` ## Performance Considerations ### System Impact - Minimal latency increase for enhanced context awareness - Efficient memory management through [Pinecone integration](https://apipie.ai/docs/features/pinecone.md) (powered by [Pinecone's vector database](https://docs.pinecone.io/docs/overview){rel=""nofollow""}) - Optimized resource utilization ### Best Practices - Use unique session IDs for different conversation contexts - Implement appropriate memory expiration times - Monitor and manage active sessions ### Resource Management - Automatic cleanup of expired sessions - Efficient handling of memory resources - Cost-effective implementation --- ## Advanced Features ### Memory Optimization - Automatic context prioritization - Intelligent memory cleanup - Resource-efficient storage ### Session Control - Fine-grained session management - Custom expiration settings - Independent memory contexts ### Integration Support - Compatible with [all supported AI models](https://apipie.ai/docs/Models/Overview.md) - Seamless API integration - Flexible implementation options By implementing Integrated Model Memory, developers can create more intelligent, context-aware AI applications while maintaining efficient resource usage and optimal performance. Learn more about enhancing your AI applications with our [Completions API](https://apipie.ai/docs/features/completions.md) and [RAG tuning capabilities](https://apipie.ai/docs/features/ragtune.md). ## Additional Resources - [Vector Similarity Search Fundamentals](https://www.pinecone.io/learn/vector-similarity/){rel=""nofollow""} - Pinecone's guide to vector search - [Anthropic's prompt caching](hhttps://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) - [OpenAI's prompt caching](https://platform.openai.com/docs/guides/prompt-caching){rel=""nofollow""} --- # Pinecone Integration: Customize Your Vector DB ::div{.docs-image-row} ![Pinecone Integration Banner](https://apipie.ai/img/docs/features/pinecone-banner.svg){width="100%"} :: Our Pinecone integration provides advanced control for users who want to manage their vector database directly. Unlike traditional RAG Tuning, which abstracts much of the complexity, the Pinecone integration allows for full control over the RAG process, including indexing and querying vectors with your own configurations. ## Why Use Pinecone Integration? Pinecone Integration gives businesses the ability to: - **Control Vector Storage**: Manage and query your vectors directly with flexible indexing and namespace options. - **Full RAG Process Control**: Instead of automatic handling, users can manage their vector collection and tuning process entirely. - **Scalability**: Perfect for larger applications that require fine-tuned control over vectors and their associated metadata. With our system, collections act as a combination of Pinecone's **index** and **namespace**, and users can specify the vector size during collection creation, making it highly customizable and scalable for any AI application. ## Creating a Vector Collection Creating a vector collection in our system combines the concepts of **index** and **namespace** from Pinecone. Here's how to create a collection: ```bash curl -L -X POST 'https://apipie.ai/v1/vectors' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: ' \ --data-raw '{ "collectionName": "my-collection", "dimension": 512 }' ``` This request creates a new vector collection with a specific dimension. You can define the size of your vectors by setting the `dimension` parameter. ## Listing Vector Collections To retrieve a list of all vector collections under your account: ```bash curl -L -X POST 'https://apipie.ai/v1/vectors/listcollections' \ -H 'Accept: application/json' \ -H 'Authorization: ' \ --data-raw '{}' ``` This returns an array of vector collections that have been created. Remember, our **collection** refers to both the index and namespace. ## Deleting a Collection or Vectors You can delete an entire collection or specific vectors within it. Here’s how to do it: ```bash curl -L -X POST 'https://apipie.ai/v1/vectors/delete' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: ' \ --data-raw '{ "collection": "my-ragtune-collection", "collectionName": "my-collection", "deleteAll": false, "ids": [ "id1", "id2" ], "filter": { "key": "value" } }' ``` Use this request to delete specific vectors by ID or delete the entire collection if needed. ## Upserting a Vector You can insert or update (upsert) vectors in your collection with metadata and embeddings: ```bash curl -L -X PUT 'https://apipie.ai/v1/vectors/upsert' \ -H 'Content-Type: application/json' \ -H 'Authorization: ' \ --data-raw '{ "collectionName": "my-collection", "vectorId": "vector-id-123", "metatag": "sampleTag", "data": "This is some clear text data associated with the vector", "embedding": [ 0.1, 0.2, 0.3 ] }' ``` This upserts a vector into the collection, allowing you to include both metadata and embeddings for future querying. ## Listing Vector IDs in a Collection To list vector IDs from a collection: ```bash curl -L -X POST 'https://apipie.ai/v1/vectors/list' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: ' \ --data-raw '{ "collectionName": "my-collection", "limit": 100, "prefix": "prefix", "paginationToken": "token123" }' ``` You can paginate through results using the `paginationToken` and set a limit on how many vector IDs to return. ## Fetching a Specific Record To fetch the contents of a specific vector by its ID: ```bash curl -L -X GET 'https://apipie.ai/v1/vectors/fetch' \ -H 'Accept: application/json' \ -H 'Authorization: ' \ --data-raw '{ "collectionName": "my-collection", "ids": "vector-id-123" }' ``` This request retrieves the specified vector's content and metadata from your collection. ## Querying Vectors You can query the vectors in your collection to find the most relevant ones: ```bash curl -L -X POST 'https://apipie.ai/v1/vectors/query' \ -H 'Content-Type: application/json' \ -H 'Authorization: ' \ --data-raw '{ "collectionName": "my-collection", "vector": [ 0.1, 0.2, 0.3 ], "topK": 10, "includeValues": true, "includeMetadata": false, "filter": { "key": "value" }, "metatag": "tag123" }' ``` This request allows you to query vectors based on their embeddings, metadata, or additional filters. ## Using Vector Collections with a Query Once you’ve created a collection and upserted vectors, you can integrate this vector collection into any query using the `rag_tune` parameter. This approach allows you to augment your AI model with your custom vector data for enhanced responses: ```bash curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "messages": [ { "role": "user", "content": "Why is the sky blue?" } ], "model": "gpt-3.5-turbo", "provider": "openai", "rag_tune": "my-collection" }' ``` This adds the vector-based augmentation from your specified collection into the AI model’s query. ## Benefits of Pinecone Integration - **Direct Control Over Vectors**: Full flexibility to create, upsert, query, and delete vectors on your terms. - **Seamless Integration**: Use any model we support with RAG via the `rag_tune` parameter, empowering users to manage custom embeddings effectively. - **Optimized Costs**: With efficient read/write units and affordable storage fees, users can manage large-scale vector databases without breaking the bank. - **Scalability**: The ability to handle high-dimensional vector spaces and large datasets, making it ideal for enterprise-level applications. ## Pinecone Integration Fees Keep in mind that Pinecone Integration incurs the following costs: - **Read/Write Operations**: Fees apply when you perform read or write operations on vector data. - **Storage Fees**: Daily storage fees accumulate based on the size of your vector collections, offering affordable rates even for large collections. ## FAQs 1. **How do I create an index in Pinecone Integration?** - You don't need to create an index separately. Our system handles this by combining the concepts of **index** and **namespace** into what we call a "collection." When you create a collection using the `collectionName` and `dimension` parameters, you effectively create both the index and namespace in one step. 2. **Can I specify different dimensions for each collection?** - Yes, when creating a collection, you can specify the vector dimension that suits your needs. This allows for flexibility based on the data you're working with. 3. **What happens if I delete a vector collection?** - Deleting a vector collection removes all the vectors and data associated with it. If you need more control, you can choose to delete specific vectors within a collection instead of the entire collection. 4. **Is there a limit on the number of vectors I can store in a collection?** - There is no hard limit on the number of vectors you can store, but keep in mind that storage costs accumulate daily based on the size of the data. The system is designed to scale with your needs. 5. **Can I use my vector collection with any model?** - Yes, once you've created and populated a vector collection, you can use the `rag_tune` parameter to augment any model we support, providing more contextually accurate responses based on your custom data. ## Conclusion Our Pinecone Integration empowers businesses to fully control their vector databases and fine-tune AI responses with custom data. This feature ensures scalability, flexibility, and affordability, allowing users to manage and query vectors directly for advanced RAG capabilities. ## Links - [Pinecone API Documentation](https://apipie.ai/docs/api/vectors){rel=""nofollow""} - [Chat Completions API](https://apipie.ai/docs/api/chatcompletions){rel=""nofollow""} - [URL Share -Upload a file, get a URL](https://apipie.ai/docs/api/upload-file){rel=""nofollow""} # Historic Usage Guide: Analyze Costs & Performance ::div{.docs-image-row} ![Historic Usage Feature Banner](https://apipie.ai/img/docs/features/usage-banner.svg){width="100%"} :: Apipie provides an advanced **Historic Usage** feature that allows users to fetch detailed information on their past API queries. This level of transparency and user awareness is unmatched, as to our knowledge, no other provider offers this level of historic usage data available via API. With the Historic Usage API, developers can access comprehensive logs of past queries, including crucial details like the provider used, latency, token count, costs, and even the source IP address. This guide outlines the available API route, the parameters for making requests, and provides detailed explanations for each field in the response. The transparency provided by Apipie enables you to manage your usage and costs efficiently. ## Fetching Query History You can fetch the history of queries sorted by timestamp in descending order using the following API route: ```bash GET https://apipie.ai/v1/queries/ ``` ### Query Parameters - **limit** (optional): Limits the number of returned queries per page. :br Example: `limit=100` - **offset** (optional): Offset used for pagination. :br Example: `offset=0` ### Example API Call Here is an example of how to fetch the last 20 queries: ```bash curl -L -X GET 'https://apipie.ai/v1/queries/?limit=20&offset=0' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' ``` ## Response Structure The response is a JSON object with a list of query records, each containing detailed information. Here is an example of the response structure: ```json { "items": [ { "id": 158939, "username": "kilroy", "timestamp": "2024-10-21T19:09:39.000Z", "app": "imageai", "provider": "openai", "route": "dall-e-3", "unit": "image", "prompt_tokens": 78, "prompt_char": 367, "response_tokens": 0, "response_char": 0, "cost": "0.0800000000", "latency_ms": 15587, "source_ip": "49.130.235.231", "chat_id": null } ... ], "limit": 20, "offset": 0 } ``` ### Explanation of Fields 1. **id**: :br The internal ID of the query. This is a unique identifier for each query. Example: `158939` 2. **username**: :br The name of the user who made the request. This will typically be the username associated with the API key used. Example: `kilroy` 3. **timestamp**: :br The exact time the query was made, in ISO format. This helps to track when the request was sent and completed. Example: `2024-10-21T19:09:39.000Z` 4. **app**: :br The name of the application or API key that initiated the query. Example: `imageai` 5. **provider**: :br The AI service provider used for the query. This could be OpenAI, Together, or any other provider supported by Apipie. Example: `openai` 6. **route**: :br The unique identifier of the AI model used within the provider. This typically refers to the model's route, such as `gpt-4o`, `dall-e-3`, or `llama-2-13b-chat`. Example: `dall-e-3` 7. **unit**: :br The unit used for cost calculations. Depending on the query type, this could be `character` or `token` for text-based queries or `image` for image generation. Example: `image` 8. **prompt\_tokens**: :br The number of tokens used in the input prompt. This is applicable for language model-based queries. Example: `78` 9. **prompt\_char**: :br The number of characters in the input prompt. This is useful for assessing the size of the input in text-based queries. Example: `367` 10. **response\_tokens**: :br The number of tokens generated in the response. This is relevant for text-based responses, but may be `0` for queries like image generation. Example: `0` 11. **response\_char**: :br The number of characters in the response. Like `response_tokens`, this is typically used for text-based responses. Example: `0` 12. **cost**: :br The total cost of the query, represented as a decimal value with up to 12 decimal places. Example: `0.0800000000` 13. **latency\_ms**: :br The total latency for the query in milliseconds. This indicates how long it took for the request to be processed. Example: `15587` 14. **source\_ip**: :br The source IP addresses from which the request originated. This field provides transparency about where the request came from and is particularly useful for tracking or auditing. Example: `49.130.235.231` 15. **chat\_id**: :br For chat-based queries, this field contains the unique identifier of the conversation or chat session as provided from the underlying provider. We keep this to track abuse. This field may be `null` for non-chat queries. Example: `null` 16. **limit**: :br The `limit` parameter that was applied to the request, indicating how many records were returned in this API call. Example: `20` 17. **offset**: :br The `offset` parameter that was applied to the request, indicating the starting point of the query history returned. Example: `0`, you could use `100`to start at 100 records prior. ## Using Historic Usage for Cost and Performance Analysis By leveraging the detailed data returned from the Historic Usage API, you can track: - **Cost Management**: Use the `cost` field to analyze and manage your API expenses. By understanding how much each query costs, you can optimize your API usage and adjust for budgetary constraints. - **Latency Insights**: Analyze the `latency_ms` field to monitor performance across various queries. If certain queries consistently show high latency, you may want to adjust your AI model or routing settings. - **Token Utilization**: With the `prompt_tokens` and `response_tokens` fields, you can track how token-heavy your requests and responses are. This can help you reduce token usage where necessary. - **Security and Auditing**: The `source_ip` field provides transparency for auditing purposes, allowing you to see where each query originated. ## Conclusion The **Historic Usage** feature of Apipie stands out by offering an unprecedented level of transparency. With detailed records of each query, users have full visibility into their API usage, allowing for better cost management, performance monitoring, and security auditing. Explore this feature today to gain more control over your API interactions. # APP Shoutout - LibreChat ![LibreChat](https://apipie.ai/img/blog/July/Librechat.svg) ## LibreChat Key Features LibreChat, an open-source ChatGPT alternative, offers users the flexibility to choose from multiple AI models and providers within a unified interface, allowing for a customizable and potentially more cost-effective alternative to traditional AI chatbot platforms. ## Integrating Multiple AI Services with LibreChat LibreChat stands out for its ability to seamlessly integrate a wide range of AI services and models, including APIpie, OpenAI, Azure, OpenRouter, Google, and Anthropic (Claude) [1](https://github.com/danny-avila/LibreChat){rel=""nofollow""} [2](https://github.com/danny-avila/LibreChat/blob/main/README.md){rel=""nofollow""}. This integration extends to both remote and local AI services, encompassing options like Ollama, Cohere, Mistral AI, and Apple MLX [1](https://github.com/danny-avila/LibreChat){rel=""nofollow""}. Users can create and save custom presets, allowing them to switch between different AI endpoints and models mid-conversation [1](https://github.com/danny-avila/LibreChat){rel=""nofollow""}. The platform supports multimodal chat capabilities, enabling users to upload and analyze images with advanced models like Claude 3 and GPT-4, as well as interact with files using various AI endpoints [1](https://github.com/danny-avila/LibreChat){rel=""nofollow""} [2](https://github.com/danny-avila/LibreChat/blob/main/README.md){rel=""nofollow""}. LibreChat's flexibility and extensive integration options make it a versatile tool for users seeking to leverage multiple AI services within a single, unified interface. ::div ![Chat](https://apipie.ai/img/docs/integrations/librechat/librechat-apipie.png) :: ## Customizing LibreChat for Corporate Use LibreChat offers several customization options for corporate environments, allowing organizations to tailor the platform to their specific needs. Companies can implement branding elements such as logos, color schemes, and background images to align the interface with their corporate identity [1](https://github.com/danny-avila/LibreChat/discussions/1207){rel=""nofollow""}. The platform supports multi-user setups with secure authentication and moderation tools, making it suitable for enterprise-wide deployment [4](https://github.com/danny-avila/LibreChat){rel=""nofollow""}. Administrators can configure system-wide custom model options through the librechat.yaml file, enabling them to enforce specific settings across the organization [2](https://github.com/danny-avila/LibreChat/issues/1617){rel=""nofollow""}. Additionally, LibreChat's open-source nature allows for further customization, including the potential to add custom endpoints and modify model specifications to meet unique corporate requirements [1](https://github.com/danny-avila/LibreChat/discussions/1207){rel=""nofollow""} [2](https://github.com/danny-avila/LibreChat/issues/1617){rel=""nofollow""}. While some features like customizing the chatbot's label and default prompts are still in development, the LibreChat community actively works on expanding corporate-friendly features to enhance its suitability for business environments [2](https://github.com/danny-avila/LibreChat/issues/1617){rel=""nofollow""} [4](https://github.com/danny-avila/LibreChat){rel=""nofollow""}. ## Cost Efficiency of Using LibreChat with APIpie.ai API LibreChat with APIpie.ai API offers a potentially cost-effective alternative to ChatGPT Plus for long-term use, especially for users with variable usage patterns. While ChatGPT Plus provides a fixed monthly rate of $20, API costs through LibreChat are usage-based, which can be more economical for some users [1](https://www.reddit.com/r/ChatGPT/comments/1clwzru/which_is_more_costeffective_for_longterm_use/){rel=""nofollow""}. However, heavy users, particularly those working on development tasks with large code inputs, may find their API costs exceeding the ChatGPT Plus subscription fee [1](https://www.reddit.com/r/ChatGPT/comments/1clwzru/which_is_more_costeffective_for_longterm_use/){rel=""nofollow""}. One user reported accruing around $0.50 in API costs for a light coding session, suggesting that regular daily usage could potentially surpass $20 per month [1](https://www.reddit.com/r/ChatGPT/comments/1clwzru/which_is_more_costeffective_for_longterm_use/){rel=""nofollow""}. Ultimately, the cost-effectiveness depends on individual usage patterns, with LibreChat offering the flexibility to optimize expenses for those who can efficiently manage their token usage. ### LibreChat - APIpie Custom Endpoint Yaml config [Here](https://www.librechat.ai/docs/configuration/librechat_yaml/ai_endpoints/apipie){rel=""nofollow""} ### APIpie.ai - LibreChat Online Demo Integration Guide [Here](https://apipie.ai/docs/Integrations/Chat-Agents/LibreChat) ### Check out the latest LibreChat v0.7.3 update :iframe{allowFullScreen="true" frameBorder="0" height="315" src="https://www.youtube.com/embed/bSVHEbVPNl4" width="560"} # Understanding AI APIs with APIpie.ai ![Understanding AI APIs](https://apipie.ai/img/blog/November/Understanding.svg) In today's rapidly advancing technological landscape, Artificial Intelligence (AI) is no longer a futuristic concept—it's a present reality transforming industries across the globe. From enhancing customer experiences to optimizing operations, AI's potential is immense. But how do businesses and developers tap into this potential without reinventing the wheel? The answer lies in **AI APIs**, especially those powered by **Generative AI**. ## What is an AI API? An **AI API (Application Programming Interface)** is a set of protocols and tools that allows developers to integrate AI functionalities into their applications seamlessly. Think of it as a bridge connecting your application to powerful AI services developed by experts. Instead of building complex AI models from scratch (a process that can be time-consuming and resource-intensive) you can leverage AI APIs to access these capabilities instantly. AI APIs enable applications to perform tasks such as: - **Natural Language Processing (NLP):** Understanding and generating human language. - **Computer Vision:** Recognizing and interpreting visual data from images and videos. - **Speech Recognition:** Converting spoken words into text. - **Predictive Analytics:** Analyzing data to forecast future trends. - **Generative AI:** Creating new content like text, images, or music based on learned patterns. ## History and Evolution of AI APIs The journey of AI APIs began with the rise of cloud computing and the need for scalable AI solutions. Initially, AI capabilities were confined to research labs and tech giants due to high computational requirements. However, as cloud services evolved, so did the accessibility of AI. - **Early 2000s:** Basic web APIs emerged, allowing for data retrieval and simple interactions. - **Mid-2010s:** Advancements in machine learning and deep learning led to more sophisticated AI models. - **Late 2010s to Present:** The advent of **Generative AI** models like GPT-3 revolutionized the field, enabling APIs to offer advanced functionalities like natural language generation and image synthesis. ## Key Features of AI APIs ### 1. **Generative Capabilities** Modern AI APIs, especially those leveraging Generative AI, can create new content based on input data. This includes generating human-like text, creating realistic images, and composing music. ### 2. **Ease of Integration** AI APIs come with comprehensive documentation and support, making it straightforward for developers to incorporate AI features into their applications using familiar programming languages. ### 3. **Scalability** Designed to handle varying workloads, AI APIs ensure consistent performance regardless of user demand, allowing your application to grow without bottlenecks. ### 4. **Cost-Effectiveness** By utilizing AI APIs, businesses save on the significant costs associated with developing and maintaining AI models and the necessary infrastructure. ### 5. **Security and Compliance** Reputable AI API providers prioritize data security, offering encryption and compliance with global standards to protect sensitive information. ## Common Use Cases for AI APIs ### **1. Content Creation with Generative AI** Businesses use Generative AI APIs to automate content creation, such as writing articles, generating product descriptions, or creating marketing materials. ### **2. Customer Service Automation** AI-powered chatbots and virtual assistants enhance customer engagement by providing instant support and personalized interactions, often using Generative AI to produce human-like responses. ### **3. Image and Video Generation** Generative AI APIs can create realistic images or modify existing ones, aiding in design, entertainment, and advertising industries. ### **4. Speech and Language Services** Transcription services convert audio to text, while translation APIs break down language barriers, enabling global communication. ### **5. Personalized Recommendations** E-commerce platforms and content providers use AI APIs to analyze user behavior and preferences, delivering tailored product or content suggestions. ## Comparing AI APIs with Other APIs While all APIs serve as intermediaries between different software applications, **AI APIs distinguish themselves through their ability to perform complex tasks that mimic human intelligence**, especially in generating new content. - **Functionality:** Traditional APIs handle data exchange and basic operations, whereas AI APIs perform advanced computations like pattern recognition and content generation. - **Complexity:** AI APIs often handle unstructured data (like images and natural language), unlike standard APIs that deal with structured data. - **Learning Capabilities:** AI APIs can improve over time through machine learning, offering enhanced performance with increased usage. ## Introducing APIpie.ai's AI API Solutions At **APIpie.ai**, our mission is to make AI accessible to businesses of all sizes. We offer a comprehensive suite of AI APIs, including **Generative AI capabilities**, designed to empower your applications with cutting-edge intelligence. ### **Why Choose APIpie.ai?** - **Generative AI Services:** Leverage our advanced Generative AI APIs to create content, generate images, or synthesize speech. - **Comprehensive AI Offerings:** From NLP and computer vision to advanced analytics, we cover all your AI needs. - **Developer-Friendly:** Our APIs are easy to integrate, with extensive documentation and sample code. - **Top-Notch Security:** We adhere to the highest security standards to protect your data. - **Exceptional Support:** Our dedicated team is here to assist you every step of the way. ## Get Started with APIpie.ai Today! Imagine transforming your application with AI-powered features in just a few lines of code. With **APIpie.ai**, this vision becomes a reality. 👉 **Ready to unlock the power of Generative AI? Visit [APIpie.ai](https://apipie.ai){rel=""nofollow""}** Join a community of innovators revolutionizing their industries with AI. Let's shape the future together. # Understanding Vector Databases in AI ![Vector Databases in AI](https://apipie.ai/img/blog/January/Understanding-Vectors.svg) Have you ever wondered how Netflix knows exactly what show to recommend next?(or at least tries to) Or how Spotify creates those perfectly curated playlists? Behind these seemingly magical recommendations lies a powerful technology: **Vector Databases**. In today's AI-driven world, these specialized databases are revolutionizing how we store, search, and understand data. ## What Makes Vector Databases Special? Imagine trying to find a specific painting in an art gallery, but instead of looking at titles or artists, you're searching for "something that feels like a sunny day at the beach." Traditional databases would struggle with such a request, but vector databases excel at understanding these kinds of semantic relationships. **Vector databases** are like highly organized art galleries that understand not just what something is, but what it means. They store data in a way that captures relationships and similarities, making it possible to find content based on meaning rather than exact matches. ### The Magic Behind Vector Databases At their core, vector databases work by converting information into numbers, lots of numbers. Here's how: 1. **Converting Data to Vectors** - Your text, images, or audio get transformed into long lists of numbers - These numbers (vectors) capture the essence of the content - Similar content gets similar number patterns 2. **Finding Similar Items** - When you search, the database looks for vectors with similar patterns - It's like finding songs that "sound similar" (Think Shazam though they use a hashing algorithm) or articles that "feel related" - This happens incredibly fast, even with millions of items ## Real-World Applications That'll Blow Your Mind ### 1. **Content Discovery** Imagine you're on YouTube watching a video about making pasta. The platform instantly suggests related cooking videos—not just based on tags or titles, but on the actual content and style of the videos. That's vector databases at work! ### 2. **E-commerce Recommendations** Ever noticed how Amazon shows you products similar to what you're viewing? Vector databases help find items that are visually or functionally similar, even if they're described differently. ### 3. **AI-Powered Knowledge Bases (RAG)** Ever noticed how ChatGPT sometimes gets facts wrong? That's where RAG (Retrieval-Augmented Generation) comes in. Vector databases help AI systems find and use relevant information from your documents, making responses more accurate and up-to-date. It's like giving AI a perfect memory of your company's knowledge! ### 4. **Fraud Detection** Banks use vector databases to spot unusual transaction patterns by comparing them with known fraud cases, protecting your money in real-time. ## The Major Players in the Vector Database World Let's meet the stars of the show: ### **[Pinecone](https://www.pinecone.io/){rel=""nofollow""}: The Enterprise Favorite** - Cloud-native and ready for serious business - Handles millions of searches per second - Perfect for production applications - [Check out our Pinecone integration](https://apipie.ai/docs/features/pinecone) ### **[Milvus](https://milvus.io/){rel=""nofollow""}: The Open Source Champion** - Free and community-driven - Supports both CPU and GPU - Great for customization ### **[Weaviate](https://weaviate.io/){rel=""nofollow""}: The Developer's Friend** - GraphQL interface makes it familiar - Built-in AI capabilities - Perfect for mixed data types ### **[Qdrant](https://qdrant.tech/){rel=""nofollow""}: The Speed Demon** - Built for performance - Flexible filtering options - Self-hosted or cloud-based ## Why Should You Care About Vector Databases? Vector databases are becoming essential because they: 1. **Make Search Smarter** - Find what you mean, not just what you type - Understand context and relationships - Work with any type of content 2. **Scale Effortlessly** - Handle millions of items - Stay fast as you grow - Adapt to your needs 3. **Enable New Possibilities** - Power recommendation systems - Enable visual search - Support AI applications ## Getting Started is Easier Than You Think With APIpie.ai's integration, you can start using vector databases in minutes. Here's a simple example: ```bash # Create a collection curl -X POST 'https://apipie.ai/v1/vectors' \ -H 'Authorization: YOUR_API_KEY' \ --data '{"collectionName": "my-first-collection"}' ``` ## The Future is Vectorized Vector databases are transforming how we interact with data. They're making applications smarter, searches more intuitive, and recommendations more accurate. As AI continues to evolve, vector databases will become even more crucial for businesses wanting to stay competitive. ### What's Next? - More sophisticated search capabilities - Better performance and efficiency - Enhanced multimodal support - Deeper AI integration ## Ready to Transform Your Applications? Vector databases are no longer just for tech giants—they're accessible to businesses of all sizes. Whether you're building a recommendation system, improving search functionality, or developing AI applications, vector databases can give you the edge you need. 👉 **Want to get started with vector databases? Visit [APIpie.ai](https://apipie.ai){rel=""nofollow""} and explore our [vector database solutions](https://apipie.ai/docs/features/pinecone).** Join the growing community of developers and businesses using vector databases to create smarter, more intuitive applications. The future of data is here—are you ready to be part of it? # Understanding RAG (Retrieval Augmented Generation) ![Understanding RAG - Retrieval Augmented Generation](https://apipie.ai/img/blog/January/Understanding-RAG.svg) Ever asked ChatGPT a question about a company's latest product, only to get a response about something from 2021? Or wondered why AI sometimes makes up information instead of using your carefully crafted documentation? Enter **[Retrieval Augmented Generation (RAG)](https://en.wikipedia.org/wiki/Retrieval-augmented_generation){rel=""nofollow""}** - the game-changing technology that's making AI responses smarter, more accurate, and actually based on your real data. ## What is RAG? Retrieval Augmented Generation (RAG) is like giving your AI a perfect memory and a research assistant. While traditional AI models rely solely on their training data, RAG actively searches through your documents to find relevant information before answering. Here's what that means: - **Retrieval**: Think of this as your AI's research phase. Before answering, it searches through your documents to find relevant information. - **Augmented**: This means "enhanced" or "improved." The AI takes what it finds and combines it with its existing knowledge. - **Generation**: Finally, it creates an answer using both its training and the retrieved information. ## How RAG Works: The Building Blocks RAG's architecture consists of four key layers working together: ### 1. Database Layer - Stores all your documents efficiently - Organizes information for quick access - Maintains your knowledge base up-to-date ### 2. Retrieval Layer - Searches through documents intelligently - Finds the most relevant information - Uses advanced matching techniques ### 3. Augmentation Layer - Combines AI knowledge with retrieved data - Enhances responses with specific information - Ensures accuracy and relevance ### 4. Network Layer - Connects all components seamlessly - Manages data flow between parts - Optimizes performance ## Why Should You Care About RAG? Imagine having a brilliant but forgetful colleague. They're incredibly smart but sometimes mix up facts or share outdated information. Now imagine giving them instant access to a company's entire knowledge base, allowing them to double-check everything before speaking. That's exactly what RAG does for AI! **RAG** combines the creative power of [large language models](https://apipie.ai/docs/Models/Overview) with the accuracy of a custom knowledge retrieval system. Instead of relying solely on what the AI learned during training (which could be outdated or irrelevant to specific needs), RAG lets it pull in specific information from available documents before generating a response. Want to dive deeper into the technical details? Check out the [original RAG paper](https://arxiv.org/abs/2005.11401){rel=""nofollow""} that started it all. ## RAG Data Management: Making Information Work Think of RAG's data handling like a highly efficient library system: ### Document Processing - Accepts various file types (PDFs, docs, spreadsheets) - Breaks down information into searchable pieces - Maintains document relationships ### Smart Retrieval - Uses semantic search to understand context - Finds information based on meaning, not just keywords - Ranks results by relevance ### Data Pipeline - Processes new documents automatically - Updates information in real-time - Maintains data freshness ## RAG Components: The Technical Side Here's what makes RAG work behind the scenes: ### Model Architecture - Embedding system for understanding text - Vector storage for efficient searching - Response generation system ### Integration Points - API connections for easy access - Monitoring tools for performance - Scaling capabilities for growth ### Optimization Features - Response time improvements - Accuracy enhancements - Resource usage management ### The "Aha!" Moment That Changes Everything Think about these frustrating AI moments we've all had: - "That's not what the product does anymore..." - "Where did it get that information from?" - "That's completely made up!" Here's how RAG fixes these headaches: - It checks your actual documents before answering - Always uses the latest information - Shows you exactly where it got its facts from - Stays focused on what you actually need to know ## How Does RAG Work Its Magic? Let's break it down: ### 1. **The Library: Your AI's Perfect Memory** Think of this as giving your AI its own research assistant: - Stores all your important documents - Handles pretty much any file type you throw at it ([check out what we support](https://apipie.ai/docs/features/ragtune)) - Keeps everything organized and searchable - [See how we process your documents](https://apipie.ai/docs/features/ragtune) ### 2. **The Smart Search: Finding What Matters** Remember our [Vector Databases blog](https://apipie.ai/blog/understanding-vector-databases-ai)? Here's where it gets cool: - Finds information lightning fast - Understands what you mean, not just what you say - Connects dots you didn't even know were there ### 3. **The Brain: Putting It All Together** Here's where the magic happens: 1. Your question comes in 2. RAG finds the perfect pieces of information 3. The AI crafts a response that's both smart AND accurate ## See RAG in Action: Real-World Magic ### **Making Customer Support Actually Helpful** Before RAG: ```text Customer: "How do I use the new feature you launched yesterday?" AI: "I don't have information about features launched after my training date." ``` After RAG: ```text Customer: "How do I use the new feature you launched yesterday?" AI: "The new Quick Export feature can be accessed by clicking the toolbar icon. Here's a step-by-step guide..." (Based on the latest documentation) ``` ### **Supercharging Your Research** - Blast through thousands of documents in seconds with our [turbocharged search](https://apipie.ai/docs/features/pinecone) - Get insights that actually make sense using smart retrieval - Connect information in ways humans might miss with semantic search ### **Creating Content That Hits the Mark** - Keep your brand's unique voice with custom settings - Never worry about accuracy with [fact-checking built in](https://apipie.ai/docs/features/integrity) - Stay on-brand with smart content filters ## What's In It For Your Business? ### 1. **Save Time (and Money!)** - Cut training time in half - Let AI handle the repetitive stuff - Get more value from what you already have ### 2. **Finally, Accuracy You Can Trust** - Say goodbye to outdated info - Get answers based on your actual data - See exactly where every fact comes from ### 3. **Scale Without the Headaches** - Make your docs work harder for you - Keep everyone on the same page - Update once, update everywhere ## The Secret Recipe for RAG Success ### 1. **Quality Matters** - Keep your docs organized (your future self will thank you) - Update regularly - Make everything clear and consistent ### 2. **Smart Searching** - Fine-tune your search with our [handy guide](https://apipie.ai/docs/features/ragtune) - Get the perfect balance of speed and accuracy with [hybrid search](https://apipie.ai/docs/features/pinecone) - Follow the pros with these [tried-and-true tips](https://www.pinecone.io/learn/retrieval-augmented-generation/){rel=""nofollow""} ### 3. **Smooth Integration** - Pick the right AI for the job - Keep an eye on performance - Stay ready for updates ## Ready to Jump In? Not sure where to start? No worries! Check out our [features](https://apipie.ai/docs/features) or see how others are making RAG work for them. Here's how quick it is to get started with APIpie.ai: ```bash # Upload your docs to a RAG collection curl -L -X POST 'https://apipie.ai/ragtune' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: ' \ --data-raw '{ "collection": "my-ragtune-collection", "url": "https://example.com/mydocument.pdf", "metatag": "important-document" }' # Let RAG do its thing curl -L -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ -H 'Authorization: Bearer ' \ --data-raw '{ "messages": [ { "role": "user", "content": "Your question here" } ], "model": "gpt-3.5-turbo", "provider": "openai", "rag_tune": "my-ragtune-collection" }' ``` ## What's Next for RAG? The future's looking bright! More businesses are discovering how RAG helps them: - Give spot-on answers every time - Scale their knowledge like never before - Keep their AI responses in check - Make their customers happier than ever ## Want to Make Your AI Smarter? Tired of your AI making things up or giving outdated answers? With APIpie.ai's RAG Tuning service, you can fix that in minutes: - Upload any kind of document - Get started right away - Only pay for what you use - Grow as big as you need 👉 **Ready to see the magic? Visit [APIpie.ai](https://apipie.ai){rel=""nofollow""} and check out our [RAG Tuning service](https://apipie.ai/docs/features/ragtune).** Join the growing crowd of businesses using RAG to make their AI actually useful. The future of AI is here—and it's a whole lot smarter with RAG! # Understanding CAG: AI's Conversation Memory ![Understanding Cache Augmented Generation (CAG)](https://apipie.ai/img/blog/March/Understanding-CAG.svg) Ever noticed how your favorite AI assistant sometimes forgets what you were just talking about? Or how you need to keep reminding it of important context from earlier in your conversation? There's a solution that's changing the game: **Cache Augmented Generation (CAG)**. Building on what we've learned about [vector databases](https://apipie.ai/docs/blog/understanding-vector-databases-ai) and [RAG systems](https://apipie.ai/docs/blog/understanding-rag-retrieval-augmented-generation-guide), CAG enhances AI responses by intelligently maintaining conversation context. ## What is Cache Augmented Generation (CAG)? Imagine if your AI could remember your entire conversation history and use that context to give you more relevant, personalized responses. That's essentially what [Cache Augmented Generation (CAG)](https://developer.ibm.com/articles/awb-llms-cache-augmented-generation/){rel=""nofollow""} does! **Cache Augmented Generation** is like giving your AI a working memory that: - Maintains a history of your conversation - Automatically includes relevant context from previous exchanges - Helps the AI understand the full context of your current question - Creates more coherent, contextually aware conversations Unlike traditional AI interactions where each question is treated in isolation, CAG ensures the AI has access to your conversation history, creating a more natural and continuous dialogue experience. ## Why CAG is a Game-Changer ### The Problem CAG Solves Let's face it - AI conversations can be frustrating when: - **Forgetful**: The AI doesn't remember what you just discussed - **Repetitive**: You have to keep providing the same context - **Disconnected**: Each response feels isolated from the conversation flow CAG tackles all these issues by maintaining conversation context across multiple interactions. ### The "Aha!" Moment Think about these common AI frustrations: - "Why do I have to keep reminding it what we're talking about?" - "I just told it that information two messages ago!" - "It's like starting over with every question!" CAG fixes these by: - Automatically including relevant conversation history - Maintaining context across multiple exchanges - Creating a coherent, flowing conversation experience ## How CAG Works Its Magic Let's break down the process: ### 1. **Conversation Memory: Beyond Single Exchanges** Traditional AI interactions treat each question in isolation. CAG is much smarter: - Stores your conversation history in a structured way - Organizes exchanges into meaningful sessions - Maintains context across multiple interactions - Uses [vector similarity search](https://www.pinecone.io/learn/vector-similarity/){rel=""nofollow""} to identify relevant past context ### 2. **Context Augmentation: Enhancing Your Current Question** When you ask a new question: - CAG analyzes what you're asking - Identifies relevant context from your conversation history - Augments your current question with this additional context - Gives the AI model a more complete picture of what you're asking This process is similar to how [APIpie's Ragtune](https://apipie.ai/docs/features/ragtune) works with documents, but applied to conversation history instead. ### 3. **Intelligent Response Generation: Better Answers** With the augmented context: - The AI understands the full conversation flow - Generates responses that acknowledge previous exchanges - Creates more coherent, contextually relevant answers - Delivers a more natural conversation experience The result is what [Google AI researchers](https://ai.googleblog.com/2020/01/towards-conversational-agent-that-can.html){rel=""nofollow""} call "conversational coherence" - the ability to maintain a consistent and natural dialogue over multiple turns. ## CAG vs. Basic Prompt Caching: What's the Difference? It's important to understand that CAG is different from simple prompt caching: ### Basic Prompt Caching (OpenAI's Approach) [OpenAI offers a simple caching system](https://platform.openai.com/docs/guides/prompt-caching){rel=""nofollow""} that: - Returns identical responses for identical prompts - Primarily focuses on efficiency and reducing duplicate processing - Doesn't enhance the context or understanding of the AI - Works only with exactly matching inputs It's like a simple lookup table - same input, same output. ### True CAG Implementation (Anthropic's Approach) [Anthropic's approach to conversation memory](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching){rel=""nofollow""} is more sophisticated: - Maintains conversation history across multiple exchanges - Intelligently selects relevant context to include - Enhances the AI's understanding of the current question - Creates more coherent, flowing conversations It's like having a conversation partner who actively remembers and references your previous exchanges. ### Side-by-Side Comparison | Feature | Basic Prompt Cache | True CAG | | ---------------------- | -------------------------- | -------------------------------------- | | Primary Purpose | Efficiency | Enhanced Context | | What It Does | Returns cached responses | Augments current question with context | | Conversation Awareness | None | High | | Implementation | Simple | More Complex | | User Experience | Faster responses | More coherent conversations | | Use Cases | Repeated identical queries | Natural flowing dialogues | ## Real-World CAG Examples That'll Make You Say "Wow!" ### **Customer Support Magic** Before CAG: ```text Customer: "I have the premium plan." AI: "Great! How can I help you with your premium plan today?" Customer: "What features do I have access to?" AI: "To tell you about available features, I'll need to know which plan you have." ``` After CAG: ```text Customer: "I have the premium plan." AI: "Great! How can I help you with your premium plan today?" Customer: "What features do I have access to?" AI: "With your premium plan, you have access to advanced analytics, priority support, and unlimited storage..." ``` ### **Personalized Assistance** - Remembers user preferences across multiple questions - Maintains context about specific projects or tasks - Creates a continuous, coherent conversation experience ### **Enhanced User Experience** Organizations implementing CAG have seen: - Reduction in users having to repeat information - Improvement in conversation coherence ratings - More natural, human-like interaction patterns ## CAG vs RAG: Short-Term Memory vs. Long-Term Knowledge Both technologies enhance AI, but they serve fundamentally different cognitive functions: ### The Human Memory Analogy Think about how your own memory works: - **Short-Term Memory (CAG/IMM)**: Remembers recent conversations and interactions. It's quick to access but limited in scope - like remembering what someone just told you a few minutes ago. - **Long-Term Memory/Reference Library (RAG)**: Stores vast amounts of knowledge accumulated over time. It takes longer to access but contains much more information - like looking up facts in an encyclopedia. CAG and RAG mirror these different memory systems: | Aspect | CAG/IMM (Short-Term Memory) | RAG (Long-Term Memory) | | ------------------ | ---------------------------------------- | --------------------------------- | | Primary Function | Remembers recent interactions | Accesses stored knowledge | | Information Source | Previous conversations | External documents/databases | | Access Speed | Extremely fast | Slightly slower (search required) | | Information Scope | Limited to past interactions | Vast knowledge repositories | | Primary Benefit | Speed & consistency | Accuracy & knowledge breadth | | Best Use Case | Repeated questions, conversation context | New information needs, research | ### Working Together Like Human Memory Just as humans use both short-term and long-term memory together, combining CAG and RAG creates a more complete AI cognitive system: - **CAG/IMM** provides the immediate context and conversation history - "What were we just talking about?" - **RAG** provides the factual knowledge and deeper information - "Let me look that up for you." This combination creates AI systems that are both responsive and knowledgeable - they remember your conversation while also being able to retrieve specific facts from their "library" when needed. ## APIpie's Integrated Model Memory (IMM): CAG Evolved At [APIpie.ai](https://apipie.ai){rel=""nofollow""}, we've taken CAG to the next level with our [Integrated Model Memory (IMM)](https://apipie.ai/docs/features/imm) system. IMM is our advanced implementation of Cache Augmented Generation that offers unique capabilities not found in other solutions: ### What Makes IMM Special - **Model-Independent Memory**: Unlike traditional CAG systems tied to specific models, IMM works seamlessly across [all supported AI models](https://apipie.ai/docs/Models/Overview) - **Cross-Model Context Retention**: Start a conversation with GPT-4, continue with Claude, and switch to Mistral while maintaining complete context - **Multi-Session Support**: Create independent memory instances for different users or applications - **Intelligent Expiration Handling**: Configure custom expiration times for conversation contexts ### How IMM Works IMM leverages our [Pinecone integration](https://apipie.ai/docs/features/pinecone) for efficient vector storage and similarity search, enabling: 1. **Automatic Context Management**: No need to manually track conversation history 2. **Seamless Conversation Tracking**: Maintain context across multiple interactions 3. **Smart Memory Controls**: Configure expiration times and clear sessions when needed 4. **Effortless Implementation**: Enable with a simple parameter in your API calls ## Getting Started with IMM: Simpler Than You Think Implementing our advanced CAG solution is surprisingly easy: ```bash # Enable Integrated Model Memory for your API calls curl -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Authorization: YOUR_API_KEY' \ -H 'Content-Type: application/json' \ --data '{ "messages": [{"role": "user", "content": "Your question here"}], "model": "gpt-4", "memory": 1, "mem_session": "user123", "mem_expire": 60 }' ``` ### Cross-Model Memory Example One of IMM's most powerful features is maintaining context across different AI models: ```bash # Start with GPT-4 curl -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Authorization: YOUR_API_KEY' \ -H 'Content-Type: application/json' \ --data '{ "memory": 1, "mem_session": "cross_model_test", "provider": "openai", "model": "gpt-4o", "messages": [{"role": "user", "content": "My favorite color is blue."}] }' # Continue with Claude, maintaining context curl -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Authorization: YOUR_API_KEY' \ -H 'Content-Type: application/json' \ --data '{ "memory": 1, "mem_session": "cross_model_test", "provider": "anthropic", "model": "claude-2", "messages": [{"role": "user", "content": "What's my favorite color?"}] }' ``` Learn more about implementing IMM in our [comprehensive documentation](https://apipie.ai/docs/features/imm). ## CAG Best Practices: Do's and Don'ts ### Do's: - Create logical session groupings for different users or topics - Implement appropriate session expiration times - [Combine with RAG](https://apipie.ai/docs/features/ragtune) for both context and knowledge - Use consistent session IDs to maintain conversation continuity - Structure conversations to build meaningful context ### Don'ts: - Don't mix unrelated conversations in the same session - Don't set overly long session retention periods - Don't rely solely on CAG for factual information (that's RAG's job) - Don't overlook privacy considerations for stored conversations - Don't neglect to clear sessions when conversations truly end ## Frequently Asked Questions About CAG ### When should I use CAG vs. basic prompt caching? Use basic prompt caching when you're focused on efficiency for identical repeated queries. Choose CAG when you want to create coherent, contextually aware conversations where the AI remembers previous exchanges. ### How does CAG improve conversation quality? CAG dramatically improves conversation quality by maintaining context across multiple exchanges. This means the AI understands references to previous messages, remembers details you've shared, and creates a more natural, flowing dialogue. ### Will CAG make my AI conversations more human-like? Absolutely! One of the key differences between human and typical AI conversations is that humans remember what was just discussed. CAG gives your AI this same capability, making interactions feel much more natural and less repetitive. ### Can I use CAG and RAG together? They're perfect companions! [RAG](https://apipie.ai/docs/features/ragtune) provides your AI with factual knowledge from documents and databases, while CAG gives it memory of the current conversation. Together, they create an AI that's both knowledgeable and contextually aware. ### What infrastructure do I need for CAG? True CAG requires vector storage capabilities and conversation management systems. With APIpie.ai's [Integrated Model Memory](https://apipie.ai/docs/features/imm), we handle all this complexity for you behind a simple API. ### How does APIpie's IMM differ from other CAG implementations? Our [Integrated Model Memory](https://apipie.ai/docs/features/imm) is model-independent, allowing you to maintain conversation context across different AI models - a capability not found in other CAG solutions. This means you can switch between models mid-conversation without losing context. ## The Future of CAG The conversation memory landscape is evolving rapidly: - More sophisticated context selection algorithms - Multi-modal conversation memory (remembering images, audio, etc.) - Personalized memory management based on user preferences - Long-term relationship building between users and AI - Integration with other AI enhancement techniques According to [recent research](https://arxiv.org/abs/2305.05176){rel=""nofollow""}, conversation memory systems like CAG will become increasingly important as users expect more natural, coherent interactions with AI systems. ## Ready to Supercharge Your AI Conversations? CAG isn't just another tech buzzword—it's a practical solution that delivers real benefits: - More coherent, flowing conversations - Reduced need for users to repeat information - More natural, human-like interactions - Better user satisfaction and engagement 👉 **Want to implement advanced conversation memory in your AI applications? Visit [APIpie.ai](https://apipie.ai){rel=""nofollow""} and explore our [Integrated Model Memory](https://apipie.ai/docs/features/imm).** Join the growing community of businesses using APIpie's [Integrated Model Memory](https://apipie.ai/docs/features/imm) to create AI experiences that truly remember what matters. The future of intelligent, contextually aware AI is here—are you ready to embrace it? # Top 5 AI Coding Models of March 2025 ![AI Coding Models](https://apipie.ai/img/blog/March/Coding.svg) The past year has brought a new generation of AI models purpose-built for coding tasks. These include: - **[OpenAI's GPT-4o](https://openai.com/index/hello-gpt-4o/){rel=""nofollow""}** (cost-optimized variant of GPT-4) - **OpenAI's "o-series" reasoning models** (often called GPT o1/o3) - **[Anthropic's Claude 3.5/3.7 "Sonnet" models](https://www.anthropic.com/news/claude-3-7-sonnet){rel=""nofollow""}** - **[DeepSeek Chat V3 & DeepSeek Reasoner R1](https://api-docs.deepseek.com/news/news250120){rel=""nofollow""}** - **[xAI's Grok v3](https://x.ai/news/grok-3){rel=""nofollow""}**, **[Meta's Llama 3](https://ai.meta.com/blog/meta-llama-3/){rel=""nofollow""}** (8B–70B), and **[Cohere's Command R+](https://docs.cohere.com/v2/docs/command-r-plus){rel=""nofollow""}** These models have been rigorously benchmarked on coding-specific tests, including [HumanEval](https://paperswithcode.com/sota/code-generation-on-humaneval){rel=""nofollow""} (programming problem-solving), MBPP (Python benchmarks), and [SWE-bench](https://github.com/princeton-nlp/SWE-bench){rel=""nofollow""} (real-world software issue resolution). All of these models are available through [APIpie's unified API](https://apipie.ai/docs/Models/Overview), making it easy to integrate them into your development workflow. ## Performance & Accuracy On major coding benchmarks, top-tier models have pushed past previous limits: - **[Claude 3.5 Sonnet](https://apipie.ai/docs/models/claude)** achieved **92%** on HumanEval, slightly edging out [GPT-4o](https://apipie.ai/docs/models/openai)'s **90.2%** - **Claude 3.7 Sonnet** scored a record-breaking **70.3% accuracy** on SWE-bench, far ahead of OpenAI's o1 (\~49%) Unlike older models that primarily generated boilerplate code, these new AI systems can debug, reason, and synthesize solutions at near-human proficiency. For more on how these capabilities are transforming development workflows, check out our article on [Understanding AI APIs](https://apipie.ai/blog/Understanding%20AI%20APIs). ## Reasoning & Debugging Modern coding AI can now analyze, debug, and fix real-world issues. SWE-bench evaluates multi-file bug fixing, and the latest results confirm a widening performance gap: - **[Claude 3.7 Sonnet](https://apipie.ai/docs/models/claude)**: **70.3% accuracy** (new record) - **OpenAI's o1/o3-mini**: \~**49% accuracy** - **[DeepSeek R1](https://apipie.ai/docs/models/deepseek)**: \~**49% accuracy** Claude 3.7's "extended reasoning" capability allows it to break down complex bugs step by step. Meanwhile, OpenAI's o-series introduces adjustable "reasoning effort" to allow deeper logical analysis. Developers note that **Claude 3.5/3.7 often provides more complete fixes, while GPT-4o is faster but may occasionally overlook subtle context issues**. ## Speed & Cost Efficiency One major 2025 trend? Faster and cheaper AI models that still perform well: - **[GPT-4o](https://apipie.ai/docs/models/openai)** was designed to be more affordable and responsive than previous GPT-4 models, making it the go-to for real-time coding assistance. - **[Claude 3.7](https://apipie.ai/docs/models/claude)**, though slower per request, often requires fewer retries, making it efficient for complex tasks. - **[Cohere Command R+](https://apipie.ai/docs/models/cohere)** is optimized for enterprise-level deployments, emphasizing low-cost, high-reliability coding output. - **OpenAI's o3-mini and o1** offer fast, low-cost options for iterative coding workflows. As AI adoption grows, many tools now mix and match models, using fast AIs for drafts and high-accuracy models for final verification. --- ## Comparison of Top AI Coding Models (March 2025) ### **[Claude 3.7 Sonnet (Anthropic)](https://apipie.ai/docs/models/claude) — The Best for Complex Debugging & Reasoning** - **💡 Accuracy:** \~92% [HumanEval](https://paperswithcode.com/sota/code-generation-on-humaneval){rel=""nofollow""}, **70.3% SWE-bench (Record high)** - **🔥 Strengths:** Best-in-class reasoning, "extended thinking" for multi-step problems, very low hallucination rate. - **📏 Context Window:** 128K+ tokens, making it ideal for handling large codebases. - **⚡ Speed & Cost:** Slower & costlier per call, but fewer retries needed, making it efficient overall. - **✅ Best For:** Large-scale debugging, complex problem-solving, and enterprise coding workflows. ### **[GPT-4o & OpenAI o-Series](https://apipie.ai/docs/models/openai) — The Workhorse for Developers** - **💡 Accuracy:** \~90% HumanEval, \~49% SWE-bench (OpenAI o1). - **🔥 Strengths:** Fastest high-accuracy model, real-time autocomplete, excellent reasoning in structured tasks. - **📏 Context Window:** 128K tokens (GPT-4o), slightly lower for mini models (o3-mini). - **⚡ Speed & Cost:** Optimized for low latency & cost, widely used in tools like [GitHub Copilot](https://github.com/features/copilot){rel=""nofollow""}. - **✅ Best For:** Everyday coding, real-time suggestions, and cost-efficient AI assistance. ### **[Google Gemini (Code-Tuned)](https://apipie.ai/docs/models/google) — Best for Large-Context Tasks** - **💡 Accuracy:** \~85%+ HumanEval (estimated) (Not publicly available for SWE-bench). - **🔥 Strengths:** Excels in contextual understanding of entire codebases, great for multi-file refactoring. - **📏 Context Window:** Up to 32K tokens (Pro version), optimized for large-scale project management. - **⚡ Speed & Cost:** Competitive speed, optimized for Google's TPU cloud deployment. - **✅ Best For:** Developers using Google Cloud, Android Studio, or those working with large repositories. ### **[Cohere Command R+](https://apipie.ai/docs/models/cohere) — The Enterprise AI Challenger** - **💡 Accuracy:** \~88% HumanEval (Unofficial), no public SWE-bench results. - **🔥 Strengths:** Optimized for [retrieval-augmented generation (RAG)](https://apipie.ai/blog/understanding-rag-retrieval-augmented-generation-guide), excellent in code search + generation tasks. - **📏 Context Window:** 16K–32K tokens, supports structured multi-step workflows. - **⚡ Speed & Cost:** Generally faster than GPT-4 on single-turn tasks, widely deployed in AWS, Azure, and Oracle AI ecosystems. - **✅ Best For:** Enterprise software teams, scalable AI integration, and structured programming tasks. ### **[DeepSeek Chat V3 & R1](https://apipie.ai/docs/models/deepseek) — The Rising Challenger** - **💡 Accuracy:** \~90% HumanEval (estimated), \~49% SWE-bench (comparable to OpenAI's o1). - **🔥 Strengths:** Blends strong coding + reasoning with an MoE (Mixture of Experts) architecture. - **📏 Context Window:** 16K tokens, well-suited for structured problem-solving. - **⚡ Speed & Cost:** More efficient than dense 70B models, moderate pricing via API access. - **✅ Best For:** Advanced developers using custom AI setups, [OpenRouter](https://openrouter.ai/){rel=""nofollow""} integrations, and experimental coding assistants. --- ## Final Thoughts The AI coding landscape is evolving rapidly, with **[Claude 3.7](https://apipie.ai/docs/models/claude) and [GPT-4o](https://apipie.ai/docs/models/openai) currently leading the pack**. However, [Google's Gemini](https://apipie.ai/docs/models/google), [Cohere Command R+](https://apipie.ai/docs/models/cohere), and [DeepSeek](https://apipie.ai/docs/models/deepseek) are closing the gap in specialized areas. Expect major advancements later in 2025 with rumored launches of **GPT-5 and Claude 4**, pushing AI coding to even greater heights. --- ### **Sources** 1. **[APIpie (AI Super Aggregator)](https://apipie.ai/dashboard){rel=""nofollow""}** 2. **[HumanEval Benchmark (Code Generation) - Papers With Code](https://paperswithcode.com/sota/code-generation-on-humaneval){rel=""nofollow""}** 3. **[Anthropic's stealth enterprise coup: How Claude 3.7 is becoming the coding agent of choice | VentureBeat](https://venturebeat.com/ai/anthropics-stealth-enterprise-coup-how-claude-3-7-is-becoming-the-coding-agent-of-choice/){rel=""nofollow""}** 4. **[OpenAI GPT-4o Benchmark - Detailed Comparison with Claude & Gemini](https://openai.com/index/hello-gpt-4o/){rel=""nofollow""}** 5. **[DeepSeek API: A Guide With Examples and Cost Calculations](https://api-docs.deepseek.com/){rel=""nofollow""}** 6. **[AWS Marketplace: Cohere Command R+ (H100) - Amazon.com](https://aws.amazon.com/marketplace/pp/prodview-jnw3ilzmy7tcg){rel=""nofollow""}** 7. **[Google Gemini Code Generation Performance](https://developers.google.com/gemini-code-assist/resources/release-notes){rel=""nofollow""}** 8. **[SWE-bench: A Benchmark for Real-World Software Engineering Tasks](https://github.com/SWE-bench/SWE-bench){rel=""nofollow""}** # Understanding Fine-Tuning vs RAG: Whats best? ![Understanding Fine-Tuning vs RAG: Choosing the Right AI Customization Strategy](https://apipie.ai/img/blog/March/Fine-Tuning-vs-RAG.svg) When it comes to customizing AI models for your specific needs, you're faced with a critical decision: should you fine-tune the model itself, or implement a Retrieval Augmented Generation (RAG) system? This choice can significantly impact your project's success, affecting everything from performance and cost to maintenance requirements and scalability. In this comprehensive guide, we'll compare these two powerful approaches to help you make the right decision for your unique use case. ## What is Fine-Tuning vs. RAG? At their core, both fine-tuning and RAG are methods to customize AI behavior, but they take fundamentally different approaches. **Fine-Tuning** involves adapting a pre-trained model by: - Further training it on your specific data - Modifying the model's internal weights and parameters - Creating a customized version of the original model Think of fine-tuning as teaching a general-purpose doctor to become a specialized surgeon—the fundamental knowledge is enhanced with specialized expertise. **Retrieval Augmented Generation (RAG)**, on the other hand, is like giving an AI model access to a specialized reference library. It involves: - Keeping the original model unchanged - Creating a knowledge base of your documents - Retrieving relevant information at query time - Augmenting the AI's prompt with this retrieved context If fine-tuning is training a specialized surgeon, RAG is giving a general doctor instant access to specialized medical textbooks exactly when needed. For a deeper dive into RAG, check out our detailed guide on [Understanding RAG: Retrieval Augmented Generation](https://apipie.ai/docs/blog/understanding-rag-retrieval-augmented-generation-guide). ## Why the Choice Matters The decision between fine-tuning and RAG isn't just a technical one—it has significant business implications: ### Cost Implications - **Fine-Tuning**: Higher upfront costs for training, potentially lower per-query costs - **RAG**: Lower setup costs, ongoing storage and retrieval costs ### Performance Considerations - **Fine-Tuning**: Can achieve higher precision for specific tasks - **RAG**: More flexible, handles new information without retraining ### Maintenance Requirements - **Fine-Tuning**: Requires periodic retraining as information changes - **RAG**: Easier to update by simply modifying the knowledge base ### Development Complexity - **Fine-Tuning**: Requires ML expertise and training infrastructure - **RAG**: Focuses more on data preparation and retrieval engineering According to [Stanford's study on LLM customization methods](https://crfm.stanford.edu/2023/03/13/alpaca.html){rel=""nofollow""}, organizations should carefully evaluate these factors based on their specific use case rather than following a one-size-fits-all approach. ## How Fine-Tuning Works Fine-tuning modifies the model itself through additional training. Here's the process: ### 1. **Data Preparation** First, you need to prepare a dataset that represents the specific knowledge or behavior you want the model to learn: - Collect examples of inputs and desired outputs - Format them according to the model's requirements - Ensure data quality and representativeness - Split into training and evaluation sets ### 2. **Training Process** The actual fine-tuning process involves: - Loading a pre-trained model as the starting point - Setting appropriate learning rates and parameters - Running additional training epochs on your custom data - Monitoring for overfitting and other training issues ### 3. **Evaluation and Deployment** After training: - Evaluate the model on held-out test data - Compare performance metrics to the original model - Deploy the fine-tuned model to your production environment - Set up monitoring for ongoing performance Fine-tuning is particularly powerful when you need the model to internalize specific patterns, styles, or domain knowledge that would be difficult to capture through prompting alone. ## How RAG Works RAG keeps the model unchanged but augments its input with relevant retrieved information: ### 1. **Knowledge Base Creation** First, you build a searchable knowledge repository: - Gather your documents, data, and knowledge sources - Process and chunk them into manageable pieces - Generate vector embeddings for each chunk - Store these in a vector database like [Pinecone](https://apipie.ai/docs/features/pinecone) ### 2. **Retrieval Process** When a user query comes in: - Convert the query to the same vector space - Search for the most relevant chunks using similarity metrics - Retrieve the top matches based on relevance scores ### 3. **Context Augmentation** Before sending to the AI model: - Combine the original query with retrieved information - Structure this combined context effectively - Send the augmented prompt to the unchanged AI model ### 4. **Response Generation** The model then: - Processes the augmented input - Generates a response informed by the retrieved context - Provides an answer grounded in your specific knowledge For a visual representation of this process, consider this simplified flow: ```text User Query → Vector Embedding → Similarity Search → Retrieve Relevant Chunks → Augment Prompt → Send to LLM → Generate Response ``` ## Real-World Example: Customer Support System Let's see how both approaches would handle implementing an AI-powered customer support system for a software company: ### The Fine-Tuning Approach ```python # Example: Fine-tuning implementation (simplified) # 1. Prepare training data (pairs of customer questions and ideal answers) training_data = [ {"role": "user", "content": "How do I reset my password?"}, {"role": "assistant", "content": "To reset your password, go to the login page and click 'Forgot Password'. Follow the email instructions to create a new password."}, # Hundreds more examples... ] # 2. Fine-tune the model response = openai.FineTuning.create( training_file="file_id_for_training_data", model="gpt-3.5-turbo", suffix="customer-support-v1" ) # 3. Use the fine-tuned model completion = openai.ChatCompletion.create( model="ft:gpt-3.5-turbo:customer-support-v1", messages=[ {"role": "user", "content": "I can't log into my account"} ] ) ``` **Results:** - The model learns patterns from support interactions - Responses match company tone and policy - Limited to knowledge available during training - Requires retraining to incorporate new products or policies ### The RAG Approach ```python # Example: RAG implementation with APIpie (simplified) # 1. Query with RAG enabled def get_support_response(query): response = requests.post( "https://apipie.ai/v1/chat/completions", headers={ "Authorization": "YOUR_API_KEY", "Content-Type": "application/json" }, json={ "messages": [{"role": "user", "content": query}], "model": "gpt-4", "rag": 1, "rag_collection": "support_documentation", "rag_depth": 3 } ) return response.json() # 2. Use the system result = get_support_response("I can't log into my account") print(result["choices"][0]["message"]["content"]) ``` **Results:** - Responses incorporate up-to-date documentation - New product information is immediately available - Knowledge base can be updated without model changes - May require more tokens per query due to context inclusion ## Decision Framework: When to Choose Each Approach To help you decide which approach is right for your use case, we've created this decision flowchart: ### Choose Fine-Tuning When: - **Task Specialization**: You need the model to excel at a specific task format - **Style Consistency**: Consistent tone, format, or brand voice is critical - **Efficiency**: You need shorter responses with less context per query - **Predictable Domain**: Your knowledge domain changes infrequently - **Training Data**: You have many high-quality examples (hundreds to thousands) - **Query Patterns**: Similar questions are asked repeatedly ### Choose RAG When: - **Knowledge Freshness**: Information updates frequently - **Factual Accuracy**: Precise, up-to-date information is critical - **Transparent Sourcing**: You need to trace responses to source documents - **Diverse Queries**: Users ask wide-ranging, unpredictable questions - **Limited Examples**: You don't have enough examples for effective fine-tuning - **Scalable Knowledge**: Your knowledge base will grow substantially over time ### Comparison Table | Factor | Fine-Tuning | RAG | | ------------------------- | -------------------------- | ----------------------- | | Setup Cost | Higher (training) | Lower (indexing) | | Ongoing Cost | Lower per query | Higher per query | | Update Ease | Requires retraining | Simple document updates | | Response Speed | Faster (no retrieval step) | Slightly slower | | Knowledge Freshness | Fixed at training time | Always current | | Specialization | High for specific tasks | Flexible across domains | | Implementation Complexity | ML expertise required | Data engineering focus | | Scaling with Knowledge | Becomes unwieldy | Scales well | ## APIpie's Approach to RAG At [APIpie.ai](https://apipie.ai){rel=""nofollow""}, we've focused on making RAG implementation as simple and effective as possible with our [RAGtune](https://apipie.ai/docs/features/ragtune) system: ### Key Features of APIpie's RAG Implementation - **Seamless Vector Database Integration**: Built-in [Pinecone integration](https://apipie.ai/docs/features/pinecone) for efficient vector storage and retrieval - **Automatic Document Processing**: Handles chunking, embedding, and storage - **Multi-Model Support**: Works with all [supported AI models](https://apipie.ai/docs/Models/Overview) - **Customizable Retrieval**: Control relevance thresholds and result counts - **Simple API Interface**: Enable RAG with just a few parameters ### Implementation Example ```bash # Process documents for your knowledge base curl -X POST 'https://apipie.ai/v1/process/document' \ -H 'Authorization: YOUR_API_KEY' \ -H 'Content-Type: application/json' \ --data '{ "collection": "product_documentation", "text": "Your document content here...", "metadata": {"source": "user_manual", "version": "2.1"} }' # Query using RAG curl -X POST 'https://apipie.ai/v1/chat/completions' \ -H 'Authorization: YOUR_API_KEY' \ -H 'Content-Type: application/json' \ --data '{ "messages": [{"role": "user", "content": "How do I configure the advanced settings?"}], "model": "gpt-4", "rag": 1, "rag_collection": "product_documentation", "rag_depth": 3 }' ``` Learn more about implementing RAG with APIpie in our [comprehensive documentation](https://apipie.ai/docs/features/ragtune). ## Hybrid Approaches: Getting the Best of Both Worlds While we've presented fine-tuning and RAG as alternatives, innovative organizations are increasingly combining these approaches: ### Sequential Hybrid: Fine-Tune Then RAG In this approach: 1. Fine-tune a model on your core domain knowledge 2. Implement RAG for up-to-date or supplementary information 3. Use the fine-tuned model as the base for RAG queries This works well when you have a stable core domain with frequently changing peripheral information. ### Selective Hybrid: Task-Based Routing Another approach is to: 1. Implement both fine-tuned models and RAG systems 2. Analyze incoming queries to determine their nature 3. Route to the appropriate system based on query type For example, product information queries go to RAG, while troubleshooting follows a fine-tuned approach. ### Implementation Considerations When implementing hybrid approaches: - Ensure clear boundaries between knowledge domains - Develop effective routing mechanisms - Monitor performance to optimize the balance - Consider the additional complexity in your architecture ## Best Practices & Tips ### Fine-Tuning Best Practices #### Do's: - Create diverse, high-quality training examples - Test the fine-tuned model extensively before deployment - Monitor for drift and performance degradation - Plan for periodic retraining cycles #### Don'ts: - Don't fine-tune on inconsistent or contradictory examples - Don't expect the model to learn information not in training data - Don't fine-tune on sensitive data without proper safeguards - Don't assume fine-tuning will fix all model limitations ### RAG Best Practices #### Do's: - Chunk documents thoughtfully for meaningful retrieval - Implement relevance thresholds to prevent irrelevant context - Update your knowledge base regularly - Use metadata to enhance retrieval precision #### Don'ts: - Don't overwhelm the context window with too many retrieved chunks - Don't neglect document preprocessing and cleaning - Don't assume perfect retrieval—implement fallbacks - Don't store sensitive information without proper access controls ## Frequently Asked Questions ### How much does fine-tuning typically cost compared to RAG? Fine-tuning has higher upfront costs (typically $500-$3,000 depending on data size and model), but potentially lower per-query costs. RAG has lower setup costs but slightly higher per-query costs due to the retrieval step and larger context windows. ### How often should I retrain my fine-tuned model? This depends on how quickly your domain changes. For stable domains, quarterly updates may be sufficient. For rapidly evolving fields, monthly retraining might be necessary. Monitor performance metrics to determine optimal retraining frequency. ### Can I use RAG with my own fine-tuned model? Yes! This hybrid approach can be very effective. Use fine-tuning for core capabilities and RAG to supplement with up-to-date information. ### How much data do I need for effective fine-tuning? While it varies by use case, most effective fine-tuning projects use at least 100-1,000 high-quality examples. More complex tasks may require several thousand examples. ### Does RAG work with all types of documents? RAG works best with text-based information that can be meaningfully chunked. It can handle PDFs, Word documents, HTML, and plain text. For images, audio, or video, additional processing steps are needed to extract textual content. ### Which approach is better for multilingual applications? RAG typically handles multilingual scenarios better, as you can include documents in multiple languages in your knowledge base. Fine-tuning for multiple languages requires substantial examples in each language. ## Ready to Get Started with Customizing Your AI? The choice between fine-tuning and RAG isn't always straightforward, but understanding the tradeoffs helps you make an informed decision: - **Fine-tuning** excels when you need specialized task performance with consistent style and have stable information. - **RAG** shines when information freshness, factual accuracy, and knowledge scalability are priorities. - **Hybrid approaches** offer flexibility for complex use cases with varying requirements. 👉 **Ready to implement RAG for your organization? Visit [APIpie.ai](https://apipie.ai){rel=""nofollow""} and explore our [RAGtune system](https://apipie.ai/docs/features/ragtune) to get started today.** Whether you choose fine-tuning, RAG, or a hybrid approach, the key is aligning your technical strategy with your specific business needs and use cases. The right choice will help you build AI systems that are not just intelligent, but truly valuable for your organization. # Top 5 Agentic AI Coding Assistants April 2025 ![Agentic AI-Powered Coding Assistants](https://apipie.ai/img/blog/April/Coding-Assistants.svg) # Top 5 Agentic AI Coding Assistants April 2025 Agentic AI coding assistants are transforming software development. Unlike traditional code completion tools, these advanced IDE-integrated agents can plan solutions, modify multiple files, run tests, and iterate on code autonomously. They act as AI co-developers that use reasoning and iterative planning to tackle complex, multi-step coding tasks with minimal human intervention. As of April 2025, numerous AI coding tools claim to be "agentic," but only a few truly deliver on this promise. After researching over 10 notable solutions, we've identified the top 5 that offer the most advanced autonomous development experience. This comparison examines their degree of autonomy, supported AI models, key capabilities, limitations, and real-world practicality. ## What Makes a Coding Assistant "Agentic"? Traditional code assistants like early versions of GitHub Copilot provided single-step suggestions in response to prompts. In contrast, agentic AI coding assistants can: - Interpret a goal in natural language - Break it into an implementation plan - Write or refactor code across multiple files - Run the code or tests - Debug errors and iterate - Complete this cycle with minimal human intervention This level of autonomy accelerates development by offloading mundane coding tasks while developers focus on high-level design. However, not all solutions are equal in their capabilities, integration options, or reliability. ## The Top 5 Agentic AI Coding Assistants ### 1. **[GitHub Copilot (Agent Mode)](https://github.com/features/copilot){rel=""nofollow""}** GitHub Copilot's agent mode (part of the Copilot X initiative) transforms the popular AI pair programmer into an "autonomous peer programmer" that can perform multi-step coding tasks on command. **Degree of Autonomy:** Highly autonomous while keeping developers in the loop for safety. It determines relevant context and files, makes code modifications, runs commands, and responds to compiler or test failures automatically. It continues this plan-act loop until it achieves the goal. **AI Model Integration:** Powered by OpenAI's models (likely GPT-4 or a specialized Codex model). It doesn't support self-hosted open-source models; it's a cloud service tied to Microsoft's Azure OpenAI. **Key Capabilities:** - Multi-file refactors and edits - Terminal command execution with approval - Test-driven development loop - Transparent "Edits" panel with undo functionality - Self-debugging when code doesn't compile **Main Limitation:** Currently in preview (VS Code Insiders only), it can be slower and more token-intensive than standard Copilot, potentially increasing costs for complex tasks. ### 2. **[Cline](https://cline.bot/){rel=""nofollow""}** Cline has quickly risen to prominence as one of the most advanced open-source agentic coding assistants. It transforms Visual Studio Code into an autonomous coding environment with a unique [Plan/Act toggle](https://www.reddit.com/r/CLine/comments/1i6poy2/cline_v32_has_a_new_planact_toggle_plan_mode/){rel=""nofollow""}. **Degree of Autonomy:** Highly autonomous with user oversight at key decision points. In Plan mode, "Cline turns into an architect that gathers information, asks clarifying questions, and designs a solution for you to review." In Act mode, it implements the plan by creating/editing files, running code/tests, and even using a browser for verification. **AI Model Integration:** Flexible and model-agnostic, supporting Anthropic Claude (3.5/3.7 "Sonnet"), OpenAI GPT-4, and Google's Gemini via providers like **[APIpie.ai](https://apipie.ai/docs/Integrations/Coding/Cline)** For a full list of supported models view the **[APIpie Dashboard](https://apipie.ai/dashboard){rel=""nofollow""}** **Key Capabilities:** - Dual Plan/Act modes for architecture and implementation - Multi-file and full-project refactoring - Terminal integration for running commands - Browser launching for UI testing - Checkpoint/rollback system for safe changes - Custom tools via MCP protocol **Main Limitation:** Requires setup of API keys/accounts for models, which can be a hurdle compared to turnkey solutions. Quality of results varies based on the model used. ### 3. **[Cursor – The AI Code Editor](https://cursor.sh/){rel=""nofollow""}** Cursor is an AI-powered code editor that has embraced agentic features to automate chunks of development workflow. It started as a streamlined code editor with built-in AI chat and has evolved to include an [Agent mode](https://docs.cursor.com/chat/agent){rel=""nofollow""} with "YOLO" fully-automatic capabilities. **Degree of Autonomy:** Designed for "minimal supervision" coding assistance. The agent can read and write files, search the codebase, run terminal commands, and perform web searches. With ["YOLO mode"](https://docs.cursor.com/chat/agent#yolo-mode){rel=""nofollow""} enabled, it can execute terminal commands without asking each time—particularly useful for test-driven development loops. **AI Model Integration:** Supports OpenAI GPT-4, Anthropic Claude, and custom API keys. It implements a ["retrieval" system](https://www.cursor.com/features#finds-context){rel=""nofollow""} to augment the model's context by indexing your codebase, mitigating token limit issues. **Key Capabilities:** - Standalone AI-first code editor - ["Instant Apply"](https://www.cursor.com/features#instant-apply){rel=""nofollow""} for direct code edits - Global codebase operations - Debugging assistance - Rules feature for constraints/preferences - MCP integration for external servers **Main Limitation:** Being a standalone editor means developers might miss plugins or behaviors from their usual setup. Some users report occasional stability issues, especially on very large projects. ### 4. [QodoAI (formerly CodiumAI)](https://www.qodoai.com/){rel=""nofollow""} QodoAI takes a unique ["quality-first" approach](https://codeparrot.ai/blogs/qodoai-code-with-an-agentic-ai){rel=""nofollow""} to AI-assisted development. It provides agentic tools focused on testing, code analysis, and automated code improvement rather than free-form coding. **Degree of Autonomy:** Autonomy focused on testing and code review. [Qodo Cover](https://codeparrot.ai/blogs/qodoai-code-with-an-agentic-ai#qodo-cover){rel=""nofollow""} (the testing agent) can analyze repositories, generate unit tests, run them, and iterate to improve coverage. [Qodo Merge](https://codeparrot.ai/blogs/qodoai-code-with-an-agentic-ai#qodo-merge){rel=""nofollow""} (the PR review agent) autonomously analyzes pull requests, generating descriptions and listing potential issues. **AI Model Integration:** Likely uses OpenAI GPT-4 for high-level reasoning with possible support for open-source models for specialized tasks. The testing agent is open source for some languages, suggesting local model support. **Key Capabilities:** - Automated test generation and improvement - PR review with issue detection - Integration with major Git providers - Support for multiple languages - Focus on code maintenance and quality **Main Limitation:** Not primarily designed for initial code generation but rather for verification and improvement of existing code. Relies on tests as guardrails, which requires good specifications or existing code to infer behavior. ### 5. [Devin AI](https://app.devin.ai/){rel=""nofollow""} Devin AI has been billed as ["the world's first AI software engineer"](https://en.wikipedia.org/wiki/Devin_AI){rel=""nofollow""}—an ambitious agent that aims to autonomously plan, code, debug, and deploy projects with minimal human input. Unlike the other tools, it's a standalone cloud-based environment. **Degree of Autonomy:** Very high autonomy—it strives to handle the entire software development cycle. After receiving a natural language task, [Devin](https://cognition.ai/){rel=""nofollow""} generates an implementation plan, writes code, runs it in a sandbox, debugs issues, and continues until completion. It can search the web for solutions and spawn sub-agents for specialized tasks. **AI Model Integration:** Uses large language models (likely Anthropic's Claude and possibly OpenAI models) in a [cloud VM that includes the code execution environment](https://qubika.com/blog/devin-ai-coding-agent/){rel=""nofollow""}. As a closed platform, users don't directly choose the model. **Key Capabilities:** - End-to-end project planning and implementation - Multi-file code generation - Code execution and testing in a VM - Web browsing for research - [Multi-agent orchestration](https://en.wikipedia.org/wiki/Devin_AI#capabilities){rel=""nofollow""} - GitHub integration **Main Limitation:** Limited availability (closed beta) and being a separate platform makes integration into existing workflows challenging. While impressive in demos, real-world testing shows it still requires significant human review for complex tasks. ## Comparison Table: Top Agentic IDEs at a Glance | Solution | Autonomy Level | AI Model Support | Key Strength | Best For | | ---------------------- | -------------- | --------------------------------- | -------------------------- | -------------------------------------------- | | GitHub Copilot (Agent) | High | OpenAI (cloud) | Integration & reliability | Everyday coding with test-driven development | | Cline | High | Pluggable (Claude, GPT-4, Gemini) | Flexibility & transparency | Complex multi-step development tasks | | Cursor | Medium-High | OpenAI, Claude, custom keys | All-in-one AI editor | Rapid prototyping & iterative development | | QodoAI | Medium | Multiple models, open-source core | Quality assurance | Test generation & code maintenance | | Devin AI | Very High | Cloud-based LLMs | End-to-end automation | Experimental autonomous development | ## Future Trends in Agentic AI Development The agentic AI coding landscape is evolving rapidly. Key trends include: - **Human-AI Collaboration Patterns:** Rather than replacing programmers, these tools are becoming collaborators. The most successful workflows treat the AI as a junior developer that produces output for human review and guidance. - **Quality and Safety Focus:** Newer agentic IDEs increasingly emphasize code quality, with agents not just coding but self-checking their work through tests and static analysis. Tools like [QodoAI](https://codeparrot.ai/blogs/qodoai-code-with-an-agentic-ai){rel=""nofollow""} and [Micro Agent](https://app.daily.dev/posts/introducing-micro-agent-an-actually-reliable-ai-coding-agent-bv6cuvzwe){rel=""nofollow""} are pioneering this approach. - **Model Improvements:** As models like GPT-5 and Claude 4 emerge with greater coding prowess and larger context windows, we'll see significant leaps in what agents can accomplish. The [NVIDIA agentic AI framework](https://blogs.nvidia.com/blog/what-is-agentic-ai/){rel=""nofollow""} shows how these orchestration layers will mature. - **Workflow Integration:** Adoption will accelerate as these tools demonstrate reliability and integrate seamlessly with established platforms and CI/CD pipelines. ## Honorable Mentions While our top 5 represent the cutting edge of agentic AI coding assistants, several other notable solutions deserve recognition: - **[Tabnine Enterprise](https://www.tabnine.com/){rel=""nofollow""}:** Known for its privacy-focused approach and on-premise deployment options, Tabnine has evolved from simple code completion to more agentic features in its enterprise offering. - **[Amazon Q Developer](https://aws.amazon.com/q/developer/){rel=""nofollow""}:** AWS's AI coding companion offers strong integration with AWS services and security-focused code suggestions. - **[JetBrains AI Assistant](https://www.jetbrains.com/ai/){rel=""nofollow""}:** Deeply integrated across all JetBrains IDEs with language-specific optimizations. - **[Replit AI](https://replit.com/ai){rel=""nofollow""}:** Combines coding assistance with Replit's cloud development environment for a seamless experience. ## Conclusion Agentic AI coding assistants are transforming software development by automating routine coding tasks while keeping developers in control of high-level decisions. GitHub Copilot Agent, Cline, Cursor, QodoAI, and Devin AI each offer unique approaches to this vision, with varying degrees of autonomy and integration. As these tools continue to evolve, they promise to dramatically increase developer productivity and code quality. The rise of autonomous coding agents stands to augment developers' capabilities in unprecedented ways, making the next era of software development an exciting one where human creativity is amplified by AI automation. # APIpie.ai: Unlock AI Solutions with Our Unified API ![Welcome](https://apipie.ai/img/blog/April/welcome.svg) ## Empowering Innovation with Seamless Integration At APIpie.ai, we are redefining the landscape of artificial intelligence applications by delivering an unprecedented aggregation of AI services. Catering to both domestic and professional users, our platform simplifies access to a diverse array of AI models from leading providers across the globe. Whether you're developing next-gen apps or enhancing existing systems, APIpie.ai equips you with the tools to deploy AI effortlessly and efficiently. ## One Subscription, Limitless Possibilities Unlock the full potential of AI with just one subscription. APIpie.ai's unique super aggregator framework offers an expansive suite of AI capabilities including language processing, vision, embeddings, image generation, voice synthesis, and code automation all accessible through a unified API. Optimize your applications by selecting services based on cost, latency, or provider, ensuring you always get the best performance at the best price. ## Designed for Developers, Built for Business Our platform is crafted with a deep understanding of developer needs. Enjoy seamless integration and reduced development overhead with APIpie.ai's streamlined API that connects you to multiple services through a single endpoint. With features like cost-effective routing, comprehensive model redundancy, and a broad selection of AI models, we empower you to focus more on creating and less on configuring. ## Join Us on the Frontier of Artificial Intelligence Dive into the future with APIpie.ai, where limitless creativity meets cutting-edge technology. Start building smarter applications today and experience the power of AI like never before. Welcome aboard! # Top 5 Open-Source ChatGPT Replacements April 2025 ![Open-Source ChatGPT Replacements](https://apipie.ai/img/blog/April/OpenSource-ChatGPT-Replacements.svg) # Top 5 Open-Source ChatGPT Replacements (April 2025) Open-source ChatGPT UI replacements have exploded in popularity, offering privacy, flexibility, and advanced agentic features that rival or surpass the official ChatGPT web app. Whether you want to self-host, integrate local LLMs, or build custom AI agents, these platforms provide robust alternatives for individuals, teams, and enterprises. This guide presents the top 5 open-source ChatGPT UI replacements with agent capabilities, based on active development, feature set, model support, and community adoption. A list of honourable mentions follows for those seeking specialized or lightweight solutions. --- ## What Makes a Great Open-Source ChatGPT Replacement? Key criteria for selection: - **Open-source license and active development** - **Self-hosted deployment (Docker, desktop, or cloud)** - **Support for multiple LLM providers ([OpenAI](https://openai.com/){rel=""nofollow""}, [Anthropic](https://www.anthropic.com/){rel=""nofollow""}, [APIpie](https://apipie.ai){rel=""nofollow""}, local models, etc.)** - **Agentic features**: tool use, function calling, RAG, plugin systems, or multi-agent orchestration - **User-friendly, modern UI with customization options** - **Community support and documentation** --- ## The Top 5 Platforms ### 1. [LibreChat](https://github.com/danny-avila/LibreChat){rel=""nofollow""} **Overview:**:br LibreChat is a polished, open-source ChatGPT clone that unifies multiple AI backends under a familiar, user-friendly interface. It supports multi-user sessions, advanced context management, and enterprise features like OAuth2 authentication and moderation. **Agent Capabilities:** - No-code Agent Builder for custom AI assistants with tools and files - Built-in agents for document Q\&A, code interpretation, and plugin-like tool use - Supports OpenAI function calling, external APIs, and ChatGPT plugins **Model Support:**:br OpenAI, Azure, Anthropic, Google, HuggingFace, local endpoints (Ollama), and more **Deployment:**:br Docker, Node.js, or Helm charts; active community and frequent updates **Best For:**:br Teams and individuals seeking a powerful, extensible ChatGPT UI with agent and plugin support --- ### 2. [AnythingLLM](https://github.com/Mintplex-Labs/anything-llm){rel=""nofollow""} **Overview:**:br AnythingLLM is an all-in-one chat application focused on chatting with your own documents and knowledge bases, while supporting general ChatGPT-style conversations. **Agent Capabilities:** - No-code Agent Builder for custom agents with tool use - Retrieval-Augmented Generation (RAG) for document Q\&A - Modular workspaces for specialized agents and multi-user management **Model Support:**:br Local runners (llama.cpp, Ollama, NeMo), OpenAI, Azure, Anthropic, Cohere, HuggingFace, OpenRouter, and more **Deployment:**:br Desktop app (Windows/Mac/Linux), Docker, or server mode; easy setup and active development **Best For:**:br Users who want to chat with custom data, run local or cloud models, and build specialized agents --- ### 3. [Open WebUI](https://github.com/open-webui/open-webui){rel=""nofollow""} **Overview:**:br Open WebUI is a self-hosted, extensible chat platform designed for offline-first operation and multi-model orchestration. **Agent Capabilities:** - Plugin “Pipelines” for web search, document retrieval, and custom Python tools - Function calling API for tool use and agent extension - Multi-model and multi-agent orchestration in a single UI **Model Support:**:br Ollama, OpenAI-compatible APIs, LMStudio, GroqCloud, Mistral, and more **Deployment:**:br Docker, pip, or Kubernetes; responsive web UI with admin controls **Best For:**:br Power users and enterprises needing advanced orchestration, plugin support, and offline/local model integration --- ### 4. [LobeChat](https://github.com/lobehub/lobe-chat){rel=""nofollow""} **Overview:**:br LobeChat is a modern, SvelteKit-based chat platform with a sleek UI, modular architecture, and strong focus on extensibility. **Agent Capabilities:** - Function call plugin system and Agent Marketplace for community plugins - Knowledge base integration (vector store memory) - Artifacts system for rich outputs (images, graphs, etc.) **Model Support:**:br OpenAI, Anthropic, Gemini, Groq, Ollama, DeepSeek, Qwen, and custom endpoints **Deployment:**:br One-click Vercel deploy, Docker, or serverless; easy customization and frequent updates **Best For:**:br Users who value aesthetics, plugin ecosystems, and multi-modal interactions --- ### 5. [Chatbot UI](https://github.com/mckaywrigley/chatbot-ui){rel=""nofollow""} **Overview:**:br Chatbot UI is a clean, open-source chat interface supporting both local and cloud-based models, with a focus on simplicity and accessibility. **Agent Capabilities:** - Sleek, intuitive UI for multi-model chat - Database-backed storage, multimodal input, and RAG support - Extensible for developers and casual users alike **Model Support:**:br ChatGPT, Claude, Gemini Pro, Ollama, DeepSeek, and more **Deployment:**:br Desktop (Windows, Mac, Linux), mobile (iOS, Android), and web **Best For:**:br Anyone seeking a straightforward, self-hosted ChatGPT alternative with broad model support --- ## Honourable Mentions - **[GPT4All (Nomic)](https://github.com/nomic-ai/gpt4all){rel=""nofollow""}:** Desktop chat client for local LLMs, fully offline, with document Q\&A. - **[Oobabooga Text Generation Web UI](https://github.com/oobabooga/text-generation-webui){rel=""nofollow""}:** Flexible web UI for running LLMs locally, with extensive extension support. - **[Hugging Face Chat UI](https://github.com/huggingface/chat-ui){rel=""nofollow""}:** SvelteKit-based UI for HuggingFace Hub models and local APIs, with plugin and function calling support. - **[Chatbox](https://github.com/Bin-Huang/chatbox){rel=""nofollow""}:** User-friendly AI chat client supporting multiple models, both online and offline. - **[Ollama UI](https://github.com/ollama-webui/ollama-webui){rel=""nofollow""}:** Minimal web UI for interacting with Ollama servers, supporting local and remote models. --- ## Conclusion The open-source ecosystem for ChatGPT UI replacements is thriving, with platforms like LibreChat, AnythingLLM, Open WebUI, LobeChat, and Chatbot UI leading the way in agentic features, model flexibility, and user experience. Whether you need advanced orchestration, document Q\&A, or a simple chat interface, these projects offer robust, privacy-friendly alternatives to proprietary solutions. As AI models and agent frameworks continue to evolve, expect even more powerful, customizable, and user-centric chat platforms to emerge—putting the future of conversational AI firmly in your hands. # Top 10 Open-Source AI Agent Frameworks of May 2025 ![Top 10 Open-Source AI Agent Frameworks](https://apipie.ai/img/blog/May/AI-Frameworks.svg) # Top 10 Open-Source AI Agent Frameworks of May 2025 AI agent frameworks have exploded in popularity as developers shift from simply calling LLMs to building autonomous systems that can reason, plan, use tools, and even collaborate with other agents. The past year has seen remarkable innovation in this space, with new frameworks emerging to address different aspects of agent development. Whether you're building a RAG-based assistant, a multi-agent research system, or enterprise AI workflows, finding the right framework is crucial. This guide compares the top open-source AI agent frameworks of 2025, analyzing their architectures, capabilities, and optimal use cases. ## What Makes a Great AI Agent Framework? Modern AI agent frameworks provide the infrastructure for language models to: - **Plan and reason** about complex tasks - **Use tools and external APIs** (function calling) - **Maintain memory** across interactions - **Manage workflows** with error handling and retry logic - **Collaborate** with other agents in multi-agent systems - **Connect to data sources** for retrieval-augmented generation Let's examine how the top frameworks deliver these capabilities. --- ## The Top 10 AI Agent Frameworks ### 1. [AG2 (formerly AutoGen)](https://github.com/ag2ai/ag2){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/AutoGen) **Language:** Python :br**Agent Style:** Multi-agent conversation framework :br**Execution Logic:** Event-driven, asynchronous :br**Memory Support:** Conversation context, extensible :br**Tool Use:** Flexible, delegated to specialized agents :br**Notable Features:** Multi-agent collaboration, tool integration, code execution :br**Why it Stands Out:** AG2 revolutionizes agent design by framing everything as a *conversation among specialized agents*. Instead of a single agent loop, you create multiple agents (e.g., an assistant, a user proxy, a coding agent) that communicate asynchronously to solve complex tasks. With AG2's event-driven architecture, agents can work concurrently rather than sequentially, reducing bottlenecks in multi-step workflows. The framework supports human-in-the-loop workflows for oversight and feedback, allowing for both autonomous operation and human guidance when needed. AG2 is particularly strong for scenarios involving collaborative problem-solving, like a dialogue between a planning agent and an execution agent, or a research task where different agents contribute specialized knowledge. Its flexible design allows for customizable agent behaviors through system messages and specialized configurations. **GitHub:** [ag2ai/ag2](https://github.com/ag2ai/ag2){rel=""nofollow""} --- ### 2. [CrewAI](https://github.com/crewAIInc/crewAI){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/CrewAI) **Language:** Python :br**Agent Style:** Role-based, collaborative teams :br**Execution Logic:** Dynamic self-organization or scripted flows :br**Memory Support:** Shared crew context, built-in memory modules :br**Tool Use:** Python functions, API integrations :br**Notable Features:** Role-based collaboration, enterprise control plane :br**Why it Stands Out:** CrewAI's greatest strength is its intuitive approach to multi-agent orchestration. By using the metaphor of a "crew" with different roles, it makes complex agent interactions approachable. You can define agents with roles like "Researcher," "Writer," and "Critic," then let them collaborate to solve tasks. CrewAI offers two modes: self-organizing crews (where agents determine their own collaboration patterns) and explicit CrewAI Flows (where you script exact interactions). This flexibility makes it suitable for both exploratory research and production applications. With 30,000+ GitHub stars and a growing community of certified developers, CrewAI has become the framework of choice for creative, multi-perspective agent systems. **GitHub:** [crewAIInc/crewAI](https://github.com/crewAIInc/crewAI){rel=""nofollow""} --- ### 3. [LangChain](https://github.com/langchain-ai/langchain){rel=""nofollow""} & [LangGraph](https://github.com/langchain-ai/langgraph){rel=""nofollow""} - [APIpie Langchain Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/Langchain) | [APIpie LangGraph Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/LangGraph) **Language:** Python (JS version available) :br**Agent Style:** Component library (LangChain) + Graph-based agent pathways (LangGraph) :br**Execution Logic:** Chains & callbacks (LangChain), Directed acyclic graph (DAG) (LangGraph) :br**Memory Support:** Multiple memory types, persistent context, checkpoint APIs :br**Tool Use:** Extensive tool library, custom tools, agent executors :br**Notable Features:** Massive ecosystem, visual graph design, checkpointing, deterministic flows :br**Why it Stands Out:** The LangChain ecosystem provides both foundational components (LangChain) and advanced orchestration capabilities (LangGraph). LangChain offers the largest ecosystem of tools, components, and integrations in the LLM space, while LangGraph extends this foundation with graph-based reasoning. Built on top of LangChain, LangGraph represents agent reasoning steps as nodes in a directed acyclic graph (DAG). This unlocks powerful patterns like parallel processing, conditional branching, and explicit error handling. Its deterministic control flow eliminates randomness in operation order, making debugging and testing more straightforward. For complex workflows where reliability matters, the LangChain+LangGraph combination stands apart. The framework supports streaming of both final answers and intermediate steps, providing transparency into the agent's thinking process. With the entire LangChain component library available as building blocks, this ecosystem offers unparalleled flexibility. Companies like Replit, Uber, and Klarna use these frameworks for production applications where certainty, auditability, and access to a rich ecosystem of components are critical. **GitHub:** [langchain-ai/langgraph](https://github.com/langchain-ai/langgraph){rel=""nofollow""} | [langchain-ai/langchain](https://github.com/langchain-ai/langchain){rel=""nofollow""} --- ### 4. [OpenAI Agents SDK](https://github.com/openai/openai-agents-python){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/OpenAI-Agents) **Language:** Python :br**Agent Style:** Function-calling, agent handoffs :br**Execution Logic:** Structured workflows, triage patterns :br**Memory Support:** Context management, result storage :br**Tool Use:** Python functions, first-class function calling :br**Notable Features:** Guardrails, agent handoffs, OpenAI integration :br**Why it Stands Out:** As OpenAI's official agent framework, this SDK provides the most seamless integration with GPT-4 and other OpenAI models. It focuses on production readiness with features like guardrails (content filtering, output validation) and standardized patterns for agent-to-agent handoffs. The SDK emphasizes a lightweight architecture centered on function calling. Instead of complex abstractions, it provides a clean runtime to manage the prompt, function schema, and callbacks, making it accessible to developers already familiar with OpenAI's API. With nearly 10,000 GitHub stars despite being released only in March 2025, the OpenAI Agents SDK is rapidly becoming the standard for production deployments leveraging OpenAI's models. **GitHub:** [openai/openai-agents-python](https://github.com/openai/openai-agents-python){rel=""nofollow""} --- ### 5. [Google Agent Development Kit (ADK)](https://github.com/google/adk-python){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/Google-ADK) **Language:** Python :br**Agent Style:** Workflow agents, LLM agents :br**Execution Logic:** Sequential, Loop, Parallel agents :br**Memory Support:** State management, persistent agents :br**Tool Use:** OpenAPI specs, Google Cloud tools, MCP :br**Notable Features:** A2A protocol, structured workflows, GCP integration :br**Why it Stands Out:** Google's ADK provides explicit constructs for complex workflow orchestration. With first-class support for Sequential, Loop, and Parallel agents, it offers unparalleled control over multi-step agent processes. ADK excels at enterprise integration, with built-in support for OpenAPI specifications as tools and secure authentication for accessing external APIs. Its integration with Google Cloud Platform makes it ideal for teams already using GCP services. As the same toolkit used internally for Google's "Gemini" AI Apps, ADK brings enterprise-grade agent capabilities to open source. Its support for the emerging Model Context Protocol (MCP) and Agent-to-Agent (A2A) protocol positions it well for future interoperability. **GitHub:** [google/adk-python](https://github.com/google/adk-python){rel=""nofollow""} --- ### 6. [Microsoft Semantic Kernel (SK)](https://github.com/microsoft/semantic-kernel){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/Semantic-Kernel) **Language:** C# and Python (multi-language SDK) :br**Agent Style:** Skills-based, plugin architecture :br**Execution Logic:** Planner with skills orchestration :br**Memory Support:** Rich abstractions (semantic, volatile, persistent) :br**Tool Use:** Native functions, OpenAPI specs, skills :br**Notable Features:** Multi-language support, enterprise integration :br**Why it Stands Out:** Semantic Kernel stands apart with its conventional programming approach to AI agents. Instead of reinventing software architecture, SK treats AI as a natural extension of existing programming paradigms, making it accessible to enterprise developers. SK's plugin architecture organizes functionality into "Skills" – reusable modules of AI or native functions that can be composed. This encourages clean separation of concerns and reusability across projects. Its planning capabilities can automatically chain these skills to accomplish complex tasks. With first-class support for C#, Python, and Java, SK caters to enterprise development teams working across multiple languages. Its integration with Azure services makes it particularly attractive for Microsoft-centric organizations. **GitHub:** [microsoft/semantic-kernel](https://github.com/microsoft/semantic-kernel){rel=""nofollow""} --- ### 7. [Hugging Face SmolAgents](https://github.com/huggingface/smolagents){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/Smolagents) **Language:** Python :br**Agent Style:** Code-generating reflexive agents :br**Execution Logic:** "Thinking in code" reflexive loop :br**Memory Support:** Limited to model context window :br**Tool Use:** Any Python library via dynamic code :br**Notable Features:** Extreme simplicity (\~1000 LOC), code generation, Hub integration :br**Why it Stands Out:** SmolAgents takes a radically different approach to agent design. Instead of complex orchestration, it creates minimal agents that "think in Python code" – literally generating and executing Python snippets to solve problems. This minimal design (just \~1000 lines of code) offers unparalleled flexibility. The agent can use any Python library by simply importing it in generated code, and security is handled through optional sandboxing via Docker or E2B. SmolAgents is perfect for rapid prototyping, experimental research, and scenarios where you want the LLM to have maximum freedom in solving problems. Its integration with the Hugging Face Hub allows sharing and reusing agents and tools. **GitHub:** [huggingface/smolagents](https://github.com/huggingface/smolagents){rel=""nofollow""} --- ### 8. [LlamaIndex](https://github.com/run-llama/llama_index){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/Llama-Index) **Language:** Python :br**Agent Style:** Data-centric, retrieval-augmented :br**Execution Logic:** Query engines, orchestration around data :br**Memory Support:** Vector stores, indices, knowledge graphs :br**Tool Use:** Data connectors, query engines as tools :br**Notable Features:** Best-in-class RAG, multi-modal indexing :br**Why it Stands Out:** LlamaIndex specializes in connecting language models to data. While other frameworks focus on actions and reasoning, LlamaIndex excels at retrieval-augmented generation (RAG) and knowledge access. Its agent architecture centers on Query Engines that intelligently route questions to appropriate data sources (indices). LlamaIndex provides rich tools for document ingestion, chunking, and indexing, making it the go-to choice for building agents that answer questions from custom data. With built-in support for various index types (vector, keyword, knowledge graph) and connectors to dozens of data sources, LlamaIndex simplifies what would otherwise be complex data engineering. **GitHub:** [run-llama/llama\_index](https://github.com/run-llama/llama_index){rel=""nofollow""} --- ### 9. [Pydantic AI](https://github.com/pydantic/pydantic-ai){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/PydanticAI) **Language:** Python :br**Agent Style:** Schema-driven, type-safe :br**Execution Logic:** Structured I/O with validation :br**Memory Support:** Custom integration options :br**Tool Use:** Schema-validated function calling :br**Notable Features:** Type safety, multi-model support, Logfire integration :br**Why it Stands Out:** Pydantic AI brings the rigor of type checking to the unpredictable world of LLMs. Created by the team behind Pydantic (used in FastAPI), it enforces output schemas that eliminate many parsing errors and ensure consistent output formats. This type-safe approach is particularly valuable for production applications where reliability is paramount. By defining strict schemas for agent outputs, Pydantic AI reduces the need for error handling and retry logic. The framework supports multiple LLM providers (OpenAI, Anthropic, Google, etc.) and integrates with Pydantic Logfire for monitoring, making it a compelling choice for enterprise applications that can't afford unexpected behavior. **GitHub:** [pydantic/pydantic-ai](https://github.com/pydantic/pydantic-ai){rel=""nofollow""} --- ### 10. [Agno (formerly Phidata)](https://github.com/agno-ai/agno){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/Agno) **Language:** Python :br**Agent Style:** Lightweight, multimodal :br**Execution Logic:** Efficient chain-of-thought :br**Memory Support:** Long-term storage, vector stores :br**Tool Use:** Python functions, API integrations :br**Notable Features:** Extreme performance, multimodal support, API generation :br**Why it Stands Out:** Agno prioritizes performance and efficiency. While other frameworks might take seconds to instantiate an agent, Agno can do it in microseconds, using 50× less memory than alternatives like LangGraph. This efficiency enables scenarios like running thousands of lightweight agents concurrently. Beyond performance, Agno stands out for its native multimodal capabilities. It seamlessly handles text, images, and audio in agent tasks without requiring complex integrations. The framework also automatically generates FastAPI routes for agents, simplifying deployment. For developers who found other frameworks too heavy, Agno offers a compelling alternative that maintains flexibility while dramatically reducing overhead. **GitHub:** [agno-ai/agno](https://github.com/agno-ai/agno){rel=""nofollow""} --- ## Honorable Mentions ### [DSPy (Stanford NLP)](https://github.com/stanfordnlp/dspy){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/DSPy) **Language:** Python :br**Key Feature:** Declarative, self-optimizing prompts :br**Why It's Notable:** DSPy treats LLM prompting as a programming discipline, not just string templates. It allows you to decompose agent reasoning into modules and optimize prompts automatically based on examples. ### [Composio](https://github.com/ComposioHQ/composio){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/Composio) **Language:** Python, TypeScript :br**Key Feature:** 250+ pre-built tool integrations :br**Why It's Notable:** Not a standalone agent framework, but a powerful complement that provides secure access to hundreds of APIs (Slack, JIRA, Google Docs, etc.) for any agent framework. ### [PixelTable](https://github.com/pixeltable/pixeltable){rel=""nofollow""} - [APIpie Integration Guide](https://apipie.ai/docs/Integrations/Agent-Frameworks/Pixeltable) **Language:** Python :br**Key Feature:** Declarative multimodal pipelines :br**Why It's Notable:** Combines data infrastructure with agent capabilities, allowing LLMs to make decisions within data processing workflows – ideal for multimodal applications. --- ## Comparing Key Framework Features | Framework | Multi-Agent | RAG Support | Execution Style | Deployment Options | Best For | | ------------------- | ----------------- | ------------------- | -------------------------- | -------------------------- | ------------------------------------------- | | **AG2** | ✅ Built-in | 🟡 Via integration | Async, event-driven | Local, cloud | Research, collaborative agents | | **CrewAI** | ✅ Core concept | 🟡 Via agent roles | Self-organizing or flows | Self-hosted, control plane | Creative tasks, multi-perspective reasoning | | **LangGraph** | ✅ Via graph nodes | ✅ Via LangChain | DAG, deterministic | Server, FastAPI | Complex workflows, auditability | | **OpenAI SDK** | ✅ Via handoffs | 🟡 Custom functions | Function calling | Any Python environment | Production OpenAI applications | | **Google ADK** | ✅ Agent Teams | ✅ Google Cloud | Sequential, Parallel, Loop | GCP, Cloud Run | Enterprise GCP deployments | | **Semantic Kernel** | ✅ Via Skills | ✅ Memory connectors | Planner, Skills | .NET, Python, Java apps | Enterprise Microsoft ecosystem | | **SmolAgents** | 🟡 Manual setup | 🟡 Custom code | Code generation | Local, Docker, E2B | Experiments, maximum flexibility | | **LlamaIndex** | 🟡 Limited | ✅ Core focus | Query Engines | Backend services | Document Q\&A, knowledge bases | | **Pydantic AI** | 🟡 Via chaining | 🟡 Custom schema | Type-validated | Python services | Production reliability | | **Agno** | ✅ Teams of Agents | ✅ Multimodal | Efficient Chain-of-Thought | FastAPI, Python apps | High-concurrency, performance-critical | --- ## Choosing the Right Framework ### For RAG Bots & Document Q\&A - **LlamaIndex** is the clear choice for sophisticated retrieval from documents, offering specialized indices and query routing - **Semantic Kernel** for enterprise knowledge bases with Azure Search integration - **OpenAI Agents SDK** paired with retrieval tools for simpler applications ### For Multi-Agent Research & Creative Systems - **AG2** for its event-driven, asynchronous multi-agent architecture - **CrewAI** for intuitive role-based collaboration and emergent behaviors - **LangGraph** when visualization and control over agent interaction is critical ### For Production Enterprise Applications - **Google ADK** for Google Cloud deployments and structured workflows - **Semantic Kernel** for Microsoft ecosystem integration - **OpenAI Agents SDK** for production-grade applications using OpenAI models - **Pydantic AI** when type safety and validation are paramount ### For Performance-Critical Applications - **Agno** for minimal overhead and high concurrency - **SmolAgents** for maximum flexibility with minimal framework - **Composio** to enhance any framework with secure tool integrations --- ## The Future of Agent Frameworks As we look ahead, several trends are emerging in the agent framework landscape: 1. **Standardization** through protocols like MCP (Model Context Protocol) and A2A (Agent-to-Agent) 2. **Security improvements** with better sandboxing and credential management 3. **Enterprise features** including monitoring, guardrails, and compliance 4. **Performance optimization** for running many lightweight agents 5. **Observability tools** for debugging complex agent behaviors While no single framework is perfect for all use cases, the diversity of approaches ensures developers can find the right tool for their specific needs. By leveraging these frameworks, developers can focus on building intelligent agents rather than reinventing core infrastructure. --- ## Conclusion The open-source AI agent ecosystem is thriving in 2025, with frameworks catering to different architectural philosophies and use cases. Whether you need a multi-agent system, robust RAG implementation, or production-grade reliability, there's a framework that fits your requirements. At [APIpie.ai](https://apipie.ai){rel=""nofollow""}, we're working to provide unified access to the models that power these frameworks, allowing you to focus on building, not managing API integrations. By standardizing access to OpenAI, Anthropic, Google, and other providers, we simplify one crucial aspect of the agent development workflow. As these frameworks continue to evolve, the barriers to building sophisticated AI agents will continue to fall, empowering developers to create increasingly capable autonomous systems. # Cursor's 'Does Not Work with Your Current Plan or API Key' Fix ![Fix Cursor IDE API Key Error](https://apipie.ai/img/blog/July/cursor.svg) # Fix Cursor's "Model XXX Does Not Work with Your Current Plan or API Key" Error If you're reading this, you've likely encountered the frustrating **"Model XXX Does not work with your current plan or API key"** error in Cursor IDE. This limitation has affected thousands of developers since Cursor restricted free account access to premium AI models. The good news? We nativley provide a unique workaround that lets you access Claude 3.5 Sonnet, GPT-4o, Gemini Pro, and other premium models **even with a free Cursor account**. ## ⚠️ Understanding the Problem [Cursor](https://cursor.com/){rel=""nofollow""} IDE's recent policy changes have created significant barriers for developers: - **Free accounts** are restricted to GPT-4.1 and "Auto" models only - **Custom API keys** don't bypass the restriction - even if you have your own OpenAI API key - **Premium models** like Claude 3.5 Sonnet, Kimi-K2-Instruct, and Gemini Pro require a Cursor Pro subscription even if you are using your own API key and custom endpoint like APIpie. ### The Frustrating Reality Even developers who configured their own API keys and custom endpoints found themselves locked out. Cursor's backend validates your account tier **before** allowing access to premium models, making custom endpoints and BYO API keys pointless in most cases except with APIpie.ai This is particularly frustrating for developers who want to use their own API credits but are blocked by Cursor's artificial restrictions. ## 🚀 The APIpie Solution: A Game-Changing Workaround We've discovered a unique solution that leverages APIpie's **API key state management** feature to completely bypass Cursor's model restrictions. This approach works because we can pre-configure which model your API key should use, regardless of what Cursor requests. ### How It Works 1. **API Key State Override**: APIpie allows you to set a "state" on your API keys that overrides any model parameter sent in requests 2. **Cursor Compatibility**: You tell Cursor to use an allowed model (like GPT-4.1) while actually receiving responses from premium models 3. **Transparent Operation**: Cursor displays GPT-4.1 but your queries are processed by Claude 3.5 Sonnet, Kimi-K2-Instruct, or any model you choose ## 📋 Step-by-Step Implementation Guide ### Step 1: Setup your APIpie account and Cursor's initial settings. 1. **Follow our Integration Guide**: Visit [Cursor Integration Guide](https://apipie.ai/docs/Integrations/Coding/Cursor) 2. **Confirm Free Account Accessibility**: After disabling all of the models except the one usable "Free Tier" model (currently gpt-4.1) verify the settings work by asking Cursor a simple quetion. 3. **Confirm Queires reach APIpie.ai**: After testing you will see query history in the [APIpie Observability Dashboard](https://apipie.ai/profile/activity){rel=""nofollow""} with the basic free model. ::div ![APIpie.ai Observability showing free tier model](https://apipie.ai/img/blog/July/Cursor-observability2.png) :: ### Step 2: Configure API Key State (The Secret Sauce) This is the crucial step that makes the workaround possible: 1. **Find your desired model**: Visit the [APIpie.ai Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} and find your desired model and note down model name and or provider. 2. **Open Chat Interface**: Start a new conversation in Cursor 3. **Set to Ask Mode**: Ensure the chat mode is set to "ask" 4. **Use Inline CLI**: Send this command as your first message: ```text :setmodel:deepinfra/Kimi-K2-Instruct ``` Replace `deepinfra/Kimi-K2-Instruct` with your preferred provider/model ie. - `anthropic/claude-3-5-sonnet-20241022` - `openai/gpt-4o` - `openrouter/gemini-2-5-pro` - `deepseek/deepseek-v3` ::div ![Set Model from Cursor](https://apipie.ai/img/blog/July/Cursor-set.png) :: 5. **Stop Pointless Cursor Messages**: After you submit the initial "set model" command Cursor may generate pointless messages feel free to stop the chat as the model will already have been changed. 6. **Configuration Lock**: When state is set, APIpie ignores any model parameters from requests and uses your pre-configured choice, any further queries on this API key will use the set model until you clear it or change it. ### Step 5: Verify It's Working 1. **Check Display**: Cursor will continue showing GPT-4.1 ::div ![Cursor still displays gpt-4.1](https://apipie.ai/img/blog/July/Cursor-query.png) :: 2. **Monitor Usage**: Visit your [APIpie Observability Dashboard](https://apipie.ai/profile/activity){rel=""nofollow""} to see the actual model being used 3. **Test Capabilities**: Open up new chats and use Cursor as you want with any model available on APIpie. ::div ![APIpie.ai Observability showing new model](https://apipie.ai/img/blog/July/Cursor-observability.png) :: ## 💡 Why This Works When Other Methods Fail ### Traditional Workarounds (Don't Work) - ❌ **Direct API Keys**: Cursor proxies direct API via its back end for prompt processing and validates the account tier / Model before processing - ❌ **Proxy Services**: Most don't offer stateful API keys requiring Models to be sent from Cursor. ### APIpie's Unique Advantage - ✅ **State Management**: Pre-configures model selection at the API key level - ✅ **Transparent Operation**: Cursor sees allowed models while you get premium features - ✅ **Cost Control**: Pay only for what you use ## 🔧 Advanced Tips ### Model Switching You can switch models anytime during your conversation: ```text :setmodel:anthropic/claude-3-5-sonnet-20241022 ``` ### Clearing Model State To return to the default model or clear the state: ```text :clearmodel ``` ## 🔍 Troubleshooting ### Common Issues and Solutions **Issue**: Still getting restricted model responses **Solution**: Verify the `:setmodel:` command was successful by checking your [APIpie usage dashboard](https://apipie.ai/profile/activity){rel=""nofollow""} **Issue**: INVALID\_PAYLOAD status: 400 - "Model is not provided" **Solution**: Double-check the model and provider name from the [APIpie dashboard](https://apipie.ai/dashboard){rel=""nofollow""} and ensure correct spelling in the format "provider/model" **Issue**: High API costs **Solution**: Use cost-effective models like DeepSeek V3 for routine tasks, reserve premium models for complex work ## 📚 Additional Resources - **Cursor Integration Guide**: [Complete setup instructions](https://apipie.ai/docs/Integrations/Coding/Cursor) - **APIpie InlineCLI Documentation**: [Inline CLI Commands](https://apipie.ai/docs/features/inlinecli) - **Usage Monitoring**: [APIpie Observability Dashboard](https://apipie.ai/profile/activity){rel=""nofollow""} - **Support**: Join our [Discord community](https://discord.gg/hs82THc9Tw){rel=""nofollow""} for help ## 🚀 Get Started Today Ready to unlock premium AI models in Cursor? Here's your action plan: 1. **Set up APIpie**: Follow our [Cursor integration guide](https://apipie.ai/docs/Integrations/Coding/Cursor) 2. **Test the workaround**: Use the `:setmodel:` command in a new chat 3. **Monitor usage**: Check your dashboard to confirm it's working 4. **Explore models**: Try different models to find your favorites With APIpie's unique API key state management, you can bypass Cursor's restrictions and access any AI model you need, all while using your own API credits and maintaining cost control. --- *This solution leverages APIpie's innovative inline CLI feature. Learn more about our [developer-focused features](https://apipie.ai/docs/features) and start building with unrestricted AI access today.* # Smart Dashboard Grouping: Compare LMMs at a Glance ![New Feature](https://apipie.ai/img/announcements/New-Feature.png) ## 📢 New Smart Grouping on the Dashboard! 🔥 We've rolled out a new smart grouping feature on the Dashboard to make your experience even smoother and more efficient! 🚀 Now, all similar models are grouped together for quick and easy comparison. This means you can effortlessly see how pricing and latency vary across different providers and model variations at a glance. 👀 For example, as you see in the screenshot, all the GPT-4o models are neatly grouped, with model variation specific text differentiated from the base model name. We provide a great side by side comparison in a way you currently can't find elsewhere. We have availability stats, the actual real cost per million tokens captured using real cost analytics and have average latency by prompt/response size, and the "time to first chunk" to ensure the best experience for streaming chat. ## Key Benefits ✅ Quickly compare performance and cost between providers and model variations :br ✅ See real-world cost analytics rather than advertised prices :br ✅ Compare latency metrics including "time to first chunk" for streaming :br ✅ Check availability stats to ensure reliable access :br ✅ Make informed decisions based on comprehensive metrics ::div{.text-center} ![Dashboard Grouping](https://apipie.ai/img/announcements/Dashboard-Grouping.png) :: We also have some more exciting features coming soon, a waterfall of new and innovative stuff coming soon. Stay tuned for updates! The APIpie Team 🥧 # DeepSeek V3.1 Released: First AI Model with Hybrid Thinking Architecture and Enhanced Agent Capabilities ![DeepSeek V3.1 Release](https://apipie.ai/img/blog/August/Deepseek-v-3-1.svg) DeepSeek has officially released **DeepSeek V3.1**, marking a significant step forward in AI model design with the introduction of the world's first **hybrid reasoning architecture**. This groundbreaking release combines thinking and non-thinking modes in a single model while delivering substantial improvements in agent capabilities and thinking efficiency. The release represents DeepSeek's vision for the **"Agent era"** - where AI models are specifically optimized for tool use, multi-step reasoning, and complex workflow automation. ## 🚀 Key Features of DeepSeek V3.1 ### Revolutionary Hybrid Reasoning Architecture DeepSeek V3.1 introduces something unprecedented in the AI space: **one model that supports both thinking and non-thinking modes**. Users can seamlessly switch between: - **Non-thinking mode** (`deepseek-chat`) - for fast, direct responses - **Thinking mode** (`deepseek-reasoner`) - for complex reasoning tasks This hybrid approach allows users to optimize for either speed or reasoning depth depending on their specific needs, all within a single model architecture. ### Dramatically Improved Thinking Efficiency According to DeepSeek's internal testing, **V3.1-Think achieves the same performance as the previous R1-0528 model while using 20-50% fewer output tokens**. This improvement translates to: - Faster response times for complex reasoning tasks - Lower API costs for thinking-intensive applications - More efficient token usage without sacrificing quality The efficiency gains are particularly notable across challenging benchmarks like AIME 2025 (87.5 vs 88.4), GPQA (81 vs 80.1), and liveCodeBench (73.3 vs 74.8), where V3.1 matches R1-0528's performance while being significantly more efficient. ### Enhanced Agent Capabilities DeepSeek V3.1 has been specifically optimized for agent workflows through targeted post-training. The improvements are evident in several key areas: **Programming Agents:** - Improved performance on SWE-bench (code fixing tasks) - Better results on Terminal-Bench (command-line environment tasks) - Enhanced multi-file code understanding and debugging **Search Agents:** - Significant improvements on multi-step reasoning tasks (browsecomp) - Better performance on expert-level questions (HLE) - Enhanced web browsing and information synthesis capabilities ## 🔧 Developer-Focused Improvements ### Strict Function Calling Support DeepSeek V3.1's API now includes **strict mode** for function calling, ensuring that outputs conform exactly to your defined JSON schemas: ```python # Strict mode ensures perfect schema compliance response = client.chat.completions.create( model="deepseek-chat", messages=[{"role": "user", "content": "Get weather data"}], tools=[weather_tool], tool_choice={"type": "function", "function": {"name": "get_weather", "strict": True}} ) ``` ### Anthropic API Compatibility For developers already using Claude-based applications, DeepSeek V3.1 now supports **Anthropic API format**, enabling easy integration with existing Claude workflows: ```python # Drop-in compatibility with Anthropic API format response = client.messages.create( model="deepseek-chat", max_tokens=1024, messages=[{"role": "user", "content": "Hello!"}] ) ``` ### Extended Context Window Both thinking and non-thinking modes now support a **128K token context window**, enabling: - Analysis of large codebases - Processing of lengthy documents - Extended conversation memory for complex workflows ## 📊 Performance Improvements According to DeepSeek's benchmarking, V3.1 shows notable improvements in agent-specific tasks: **Programming Tasks:** - Better performance on SWE-bench (real-world code fixing) - Improved Terminal-Bench results (CLI environment tasks) **Search and Reasoning:** - Enhanced performance on browsecomp (multi-step search reasoning) - Better results on HLE (expert-level multi-disciplinary questions) The model also demonstrates improved output length control in non-thinking mode, producing more concise responses while maintaining quality compared to previous versions. ## 🔧 Technical Architecture ### Model Improvements - **UE8M0 FP8 Scale** parameter precision for improved efficiency - **Redesigned tokenizer** and chat template (incompatible with DeepSeek V3) - **840B additional training tokens** beyond the base V3 model - **Post-training optimization** specifically for tool usage and agent workflows ### Open Source Availability DeepSeek V3.1 is fully open source and available on multiple platforms: **Base Model:** - [Hugging Face](https://huggingface.co/deepseek-ai/DeepSeek-V3.1-Base){rel=""nofollow""} - [ModelScope](https://modelscope.cn/models/deepseek-ai/DeepSeek-V3.1-Base){rel=""nofollow""} **Chat Model:** - [Hugging Face](https://huggingface.co/deepseek-ai/DeepSeek-V3.1){rel=""nofollow""} - [ModelScope](https://modelscope.cn/models/deepseek-ai/DeepSeek-V3.1){rel=""nofollow""} ## 🚀 Getting Started with DeepSeek V3.1 ### Via APIpie's Unified API DeepSeek V3.1 is available now through APIpie's unified API platform: ```python import openai client = openai.OpenAI( base_url="https://apipie.ai/v1", api_key="your-apipie-key" ) # Non-thinking mode for fast responses response = client.chat.completions.create( model="deepseek-chat", provider="deepseek", messages=[{"role": "user", "content": "Explain machine learning"}] ) # Thinking mode for complex reasoning response = client.chat.completions.create( model="deepseek-reasoner", provider="deepseek", messages=[{"role": "user", "content": "Solve this complex problem step by step"}] ) ``` ### Direct DeepSeek API You can also access DeepSeek V3.1 directly through DeepSeek's OpenAI-compatible API: ```python from openai import OpenAI client = OpenAI( api_key="your-deepseek-api-key", base_url="https://api.deepseek.com" ) response = client.chat.completions.create( model="deepseek-chat", # or "deepseek-reasoner" messages=[{"role": "user", "content": "Hello!"}] ) ``` ## 💰 Pricing Updates DeepSeek has announced pricing changes effective **September 6, 2025**: - New pricing structure will take effect - Night-time discount rates will be discontinued - Current pricing remains in effect until the transition date APIpie users benefit from stable, competitive pricing regardless of these changes. ## 🎯 Who Should Use DeepSeek V3.1 ### Ideal for Agent Builders If you're building AI agents that need to: - Perform complex coding tasks - Handle multi-step search and reasoning - Use tools and function calling reliably - Process large amounts of context ### Perfect for Cost-Conscious Developers The hybrid architecture means you can: - Use fast mode for simple tasks to save costs - Switch to thinking mode only when deep reasoning is needed - Optimize your token usage without sacrificing capability ### Great for Claude Users With Anthropic API compatibility, you can: - Easily migrate existing Claude-based applications - Test DeepSeek V3.1 as a drop-in replacement - Compare performance without rewriting code ## 🔮 Looking Forward DeepSeek V3.1's hybrid architecture represents a significant evolution in AI model design, prioritizing practical utility for real-world applications. The focus on agent capabilities, improved efficiency, and developer experience signals a shift toward more practical, production-ready AI systems. The open-source availability of both base and chat models also ensures that developers have full access to experiment, fine-tune, and deploy DeepSeek V3.1 according to their specific needs. ## 🚀 Get Started Today DeepSeek V3.1 is available now through [APIpie's unified API platform](https://apipie.ai/docs/models/deepseek). Whether you're building your first AI agent or looking to upgrade existing applications, V3.1's combination of efficiency, capability, and flexibility makes it an compelling choice. Visit our [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} to start experimenting with DeepSeek V3.1 alongside 200+ other AI models, all through one simple API. --- *Stay updated with the latest AI model releases and capabilities through APIpie - your gateway to the full spectrum of AI models.* # Meet Integrated Model Memory: Cross-Model Caching ![New Feature](https://apipie.ai/img/announcements/New-Feature.png) ## Introducing Integrated Model Memory (IMM) - Now in Beta! 🧠✨ We're excited to announce **Integrated Model Memory (IMM)** – a powerful, plug-and-play memory solution that seamlessly integrates across all supported AI models! With just a simple parameter, developers can now enable persistent memory across sessions and models, eliminating the need for complex memory management. ## Key Benefits - ✅ Works across 300+ models - ✅ No extra setup—just enable memory! - ✅ Persistent context retention across conversations - ✅ Multi-user session support ## Quick Start Guide - 📖 [Feature details & usage: IMM Feature Docs](https://apipie.ai/docs/features/imm) - 🛠️ [API Reference: IMM API Docs](https://apipie.ai/docs/api/chatcompletions) ## What is IMM? IMM is our implementation of Cache Augmented Generation (CAG), but unlike traditional CAG systems, IMM works across all models! You can start a conversation with GPT-4, switch to Claude, and finish with Mistral, all while maintaining full context. ## Key Features - **Easy Implementation** – Just add `"memory": 1` to your API calls! - **Advanced Session Management** – Isolated memory for different users or use cases - **Smart Memory Controls** – Set expiration times, manage memory efficiently - **Cross-Model Context Retention** – Seamless transition between AI models - **Developer-Friendly** – No vector DB needed, fully managed memory IMM remembers your past conversations—no need to re-send context! ## Cross-Model Memory in Action 1. Start with GPT-4 2. Continue with Claude 3. Switch to Mistral Your conversation context remains intact across all models! IMM ensures full session continuity even when switching providers. ## Beta Now Live – Help Us Improve! Please report bugs so we can refine and improve IMM. Happy building, The APIpie Team 🎉 # New Billing System: Auto Top-Up & More ![New Billing System](https://apipie.ai/img/announcements/Announcement.png) ## Completely Rebuilt Billing System 💳 Big news: we've completely rebuilt our billing backend! 🎉 ## What's New ### Email Mismatch & Other Billing Fixes ✅ No more mix-ups with email addresses. Everything now matches perfectly! ### Auto Top-Up 🚀 Your credits can now be topped up automatically, so you'll always have enough to keep things running smoothly. If you spot any issues with payments 💸, just give us a shout and we'll sort it out ASAP. ## Additional Improvements This release also includes a number of other small UI bug fixes and improvements that have been pushed out along with this update. --- It's been a long time coming and hopefully we've caught all the potential issues in our QA. If you encounter anything unusual, either raise it here or ping me directly and we will get on it. Thanks for sticking with us, and have an awesome day! 😄 Cheers, The APIpie Team 🥧 # Claude 3.7 Sonnet: Hybrid Reasoning Model Now Live on APIpie ![Claude 3.7 Sonnet](https://apipie.ai/img/announcements/New-Model.png) ## 🚀 **Now Live on APIpie.ai: [Claude 3.7 Sonnet](https://www.anthropic.com/news/claude-3-7-sonnet){rel=""nofollow""}** We're thrilled to announce that Anthropic's most intelligent model to date — Claude 3.7 Sonnet — is now available through [APIpie.ai](https://apipie.ai){rel=""nofollow""}! ### 🧠 **Claude 3.7 Sonnet: The First Hybrid Reasoning Model** > • **Dual-mode operation**: Near-instant responses OR extended step-by-step thinking :br > • **Fine-grained control**: Set thinking token budgets up to 128K tokens via API :br > • **State-of-the-art coding**: Best-in-class performance on SWE-bench Verified and TAU-bench :br > • **Real-world focus**: Optimized for practical business tasks, not just competition problems ### ⚡ **Key Capabilities** > • **Superior coding performance**: Exceptional at handling complex codebases and full-stack development :br > • **Extended thinking mode**: Visible step-by-step reasoning for complex problems :br > • **Multimodal excellence**: Advanced capabilities across text, code, and visual tasks :br > • **Reduced refusals**: 45% fewer unnecessary rejections compared to predecessors ### 🛠️ **Claude Code: Agentic Development Tool** > • **Active collaboration**: Search, read, edit, test, and commit code autonomously :br > • **GitHub integration**: Direct repository connections for seamless workflow :br > • **Command line native**: Terminal-based tool for advanced development tasks :br > • **Test-driven development**: Complete complex refactoring tasks in single passes ### 🎯 **What Makes Claude 3.7 Special** > • **Unified architecture**: Single model that adapts from quick responses to deep reasoning :br > • **Industry validation**: Preferred by Cursor, Cognition, Vercel, Replit, and Canva :br > • **Production-ready**: Superior design taste and drastically reduced errors :br > • **Flexible reasoning**: Trade off speed and cost for answer quality as needed ### 💰 **Pricing** > • **Same cost as predecessors**: $3 per million input tokens, $15 per million output tokens :br > • **Thinking tokens included**: No additional cost for extended reasoning mode :br > • **Cost-effective**: Pay only for the thinking depth you need ### 📊 **Benchmark Performance** > • **SWE-bench Verified**: State-of-the-art performance on real-world software issues :br > • **TAU-bench**: Leading results on complex real-world tasks with tool interactions :br > • **Coding excellence**: Validated by top development platforms and frameworks :br > • **Instruction following**: Superior performance across general reasoning tasks ### 🔄 **Extended Thinking Mode** > • **Visible reasoning**: Watch Claude think through complex problems step-by-step :br > • **Controllable depth**: Set maximum thinking tokens for optimal cost/quality balance :br > • **Enhanced accuracy**: Significant improvements in math, physics, and coding tasks :br > • **Self-reflection**: Model evaluates its own reasoning before providing answers ### 🛡️ **Safety & Reliability** > • **Extensive testing**: Evaluated by external security and safety experts :br > • **Responsible scaling**: Comprehensive system card with detailed safety evaluations :br > • **Prompt injection resistance**: Enhanced training to resist and mitigate attacks :br > • **Nuanced understanding**: Better distinction between harmful and benign requests ### 🔁 **Next Up** > • Enhanced Claude Code capabilities with improved tool call reliability :br > • Long-running command support and better in-app rendering :br > • Expanded reasoning model features and performance optimizations :br > • Stay tuned for continued improvements to the APIpie integration! --- **Ready to experience the future of AI reasoning?** [Get started with Claude 3.7 Sonnet on APIpie today](https://apipie.ai){rel=""nofollow""} and discover how hybrid reasoning can transform your development workflow. # GPT-4.5: OpenAI's Strongest Model for Chat ![GPT-4.5](https://apipie.ai/img/announcements/New-Model.png) ## 🧠 **OpenAI Launches GPT-4.5** OpenAI has released **GPT-4.5**, their largest and best model for chat. This breakthrough in scaling unsupervised learning delivers enhanced world knowledge, improved emotional intelligence, and significantly reduced hallucinations compared to GPT-4o. 🚀 **Key Features of GPT-4.5** - **Scaling Unsupervised Learning**:br → Advanced world model accuracy and intuition :br → Broader knowledge base through scaled pre-training :br → Enhanced pattern recognition and creative insights - **Enhanced Emotional Intelligence**:br → Greater "EQ" and natural conversation abilities :br → Better understanding of human intent and subtle cues :br → More empathetic and contextually appropriate responses - **Significantly Reduced Hallucinations**:br → **37.1% hallucination rate** vs GPT-4o's 61.8% :br → **62.5% accuracy on SimpleQA** vs GPT-4o's 38.2% :br → Improved factual accuracy across knowledge domains - **Superior Creativity**:br → Enhanced aesthetic intuition and creative problem-solving :br → Better understanding of design and visual appeal :br → Improved writing assistance and content creation - **Complementary to Reasoning Models**:br → Serves as stronger foundation for reasoning agents :br → Works alongside OpenAI o1/o3-mini reasoning models :br → Future models will combine both paradigms 📡 **Performance & Capabilities** - **Academic Excellence**:br → **GPQA (Science)**: 71.4% vs GPT-4o's 53.6% :br → **AIME '24 (Math)**: 36.7% vs GPT-4o's 9.3% :br → **MMMLU (Multilingual)**: 85.1% vs GPT-4o's 81.5% :br → **MMMU (Multimodal)**: 74.4% vs GPT-4o's 69.1% - **Human Preference**:br → **56.8-63.2% win rate** vs GPT-4o across query types :br → Preferred for creative intelligence and professional queries :br → More natural and intuitive interactions - **Advanced Capabilities**:br → Superior help with writing and programming :br → Enhanced practical problem-solving abilities :br → Better aesthetic judgment for design work :br → Improved conversational warmth and empathy - **Research Preview Status**:br → More expensive than GPT-4o due to large scale :br → Long-term availability depends on community feedback :br → Currently lacks multimodal features (Voice Mode, video) 🛠️ **Available now** — plug in, route smart, and start building see it on our 👉 [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} **Learn more:** [OpenAI GPT-4.5 Introduction](https://openai.com/index/introducing-gpt-4-5/){rel=""nofollow""} # QwQ-32B: Reinforcement Learning Reasoning Model ![QwQ-32B](https://apipie.ai/img/announcements/New-Model.png) ## 🧠 **Alibaba Cloud Launches QwQ-32B** Alibaba Cloud has released **QwQ-32B**, a reinforcement learning reasoning model that achieves performance comparable to DeepSeek-R1 (671B parameters) while using only 32 billion parameters. This breakthrough demonstrates the power of scaled reinforcement learning for mathematical and coding tasks. 🚀 **Key Features of QwQ-32B** - **Reinforcement Learning Powered**:br → Trained with outcome-based rewards and accuracy verifiers :br → Multi-stage RL training for math, coding, and general capabilities :br → Cold-start approach with continuous performance improvement - **Exceptional Efficiency**:br → **32B parameters** matching 671B parameter model performance :br → Significantly lower computational requirements :br → Cost-effective deployment and inference - **Superior Reasoning Capabilities**:br → Advanced mathematical problem-solving :br → Strong coding proficiency and code generation :br → Agent functionality with tool use and environmental feedback - **Open Source Advantage**:br → **Apache 2.0 license** for maximum flexibility :br → Available on Hugging Face and ModelScope :br → No licensing fees for commercial deployment - **Agent Integration**:br → Critical thinking while utilizing tools :br → Adapts reasoning based on environmental feedback :br → Foundation for long-horizon reasoning applications 📡 **Performance & Benchmarks** - **Mathematical Reasoning**:br → State-of-the-art performance on complex math problems :br → Excellent results across various mathematical benchmarks - **Coding Excellence**:br → Superior programming capabilities :br → Advanced code generation and analysis - **General Problem-Solving**:br → Robust performance across diverse reasoning tasks :br → Strong instruction following and alignment 🛠️ **Available now** — plug in, route smart, and start building see it on our 👉 [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} **Learn more:** [QwQ-32B Official Blog](https://qwenlm.github.io/blog/qwq-32b/){rel=""nofollow""} # Gemini 2.5: Google's Most Intelligent AI Model ![Gemini 2.5](https://apipie.ai/img/announcements/New-Model.png) ## 🧠 **Google Launches Gemini 2.5** Google has released **Gemini 2.5 Pro** and **Gemini 2.5 Flash**, thinking models that achieve #1 on LMArena leaderboard by significant margin. These models excel at reasoning through complex problems before responding, with state-of-the-art coding performance and advanced multimodal capabilities. 🚀 **Key Features of Gemini 2.5** - **Advanced Thinking Models**:br → Think through complex problems before responding :br → Enhanced reasoning capabilities with step-by-step analysis :br → State-of-the-art performance across reasoning benchmarks - **World-Leading Coding**:br → **#1 on WebDev Arena** with ELO score of 1415 :br → 63.8% on SWE-Bench Verified with custom agent setup :br → Superior performance on programming tasks - **Gemini 2.5 Pro: Maximum Intelligence**:br → Top across all LMArena leaderboards for human preference :br → 18.8% on Humanity's Last Exam without tool use :br → 1 million token context window (2 million coming soon) - **Gemini 2.5 Flash: Efficient Excellence**:br → 20-30% more efficient while improving performance :br → Speed optimized for high-throughput applications :br → Generally available for production deployment - **Deep Think (Experimental)**:br → Enhanced reasoning mode considering multiple hypotheses :br → Impressive scores on 2025 USAMO benchmark :br → 84.0% on MMMU multimodal reasoning 📡 **Performance & Capabilities** - **Reasoning Excellence**:br → Leading performance on math and science benchmarks :br → State-of-the-art on GPQA and AIME 2025 :br → Superior long context and video understanding - **Advanced Features**:br → **Native audio output** with tone control and 24+ languages :br → **Computer use** capabilities through Project Mariner :br → **Enhanced security** against prompt injection attacks :br → **Live API** with audio-visual input processing - **Multimodal Capabilities**:br → Text, audio, images, video, and code repositories :br → Advanced vision-language comprehension :br → Real-time dialogue with emotion detection - **Developer Experience**:br → **Thought summaries** for structured reasoning processes :br → **Thinking budgets** to control token usage and latency :br → **MCP support** for open-source tool integration 🛠️ **Available now** — plug in, route smart, and start building see it on our 👉 [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} **Learn more:** [Google I/O 2025 Updates](https://blog.google/technology/google-deepmind/google-gemini-updates-io-2025/){rel=""nofollow""} # Llama 4: Scout and Maverick Models Now Available ![Llama 4](https://apipie.ai/img/announcements/New-Model.png) ## 🚨 **Big drop: Llama 4 now on [APIpie.ai](https://apipie.ai){rel=""nofollow""}** 🚨 We’ve just integrated the latest *Llama 4 Scout* and *Maverick* models — blazing-fast, multimodal, and context-stacked. 🎯 **What’s New?** 🌟 **Llama 4 Scout** > 🔹 109B parameters (16 experts active) :br > 🔹 10M token context window :br > 🔹 Tops Gemma 3, Mistral 3.1, and Gemini Flash-Lite 🚀 **Llama 4 Maverick** > 🔹 400B parameters (128 experts active) :br > 🔹 1M token context :br > 🔹 Beats GPT-4o & Gemini Flash in vision + code :br > 🔹 Matches DeepSeek V3 in reasoning 🧠 **Llama 4 Behemoth** *(coming soon)* > 🔹 2T parameter distillation beast :br > 🔹 Outperforms GPT-4.5 & Claude Sonnet 3.7 in STEM :br > 🔹 Still cooking... 🛠️ **Available now** — plug in, route smart, and start building see it on our 👉 [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} # Enhanced Error Handling, & Diagnostics ![Improved Error Handling](https://apipie.ai/img/announcements/Announcement.png) ## Improved Error Handling System 🛠️ We're excited to announce significant improvements to our error handling system! These enhancements deliver clearer diagnostics and a more developer-friendly experience when troubleshooting API issues. ## What's New ### 🔍 More Detailed Error Messages - Error responses now include specific details about what went wrong - Clear suggestions for how to resolve common issues - References to relevant documentation where applicable ### 📊 Better Error Categorization - Consistent HTTP status codes aligned with RESTful best practices - Logical grouping of errors by type (authentication, validation, rate limits, etc.) - Standardized error formats across all endpoints ### 🚦 Improved Error Transparency - Clear distinction between user errors and system issues - Better visibility into rate limiting with headers showing remaining quota - More predictable error behavior for easier integration ## Example Response Here's an example of our new error response format: ```json { "error": { "code": "rate_limit_exceeded", "message": "You have exceeded your current quota, please check your plan and billing details.", "param": null, "type": "quota", "details": { "limit": 60, "remaining": 0, "reset": 1703012488 }, "documentation_url": "https://apipie.ai/docs/" } } ``` ## Benefits for Developers - **Faster Debugging**: Pinpoint issues more quickly with specific error details - **Easier Integration**: Consistent error formats make error handling more predictable - **Better User Experience**: Provide more informative feedback to your end users - **Reduced Support Needs**: Clearer error messages mean fewer support tickets ## Implementation Details This update has been applied across all API endpoints and services. No changes to your integration are required to benefit from these improvements. We're committed to continuously enhancing the developer experience, and these error handling improvements represent an important step toward making our API more robust and user-friendly. As always, we welcome your feedback on these changes! The APIpie Team 🥧 # Dashboard V2: Global AI Operations Center ![Dashboard V2](https://apipie.ai/img/announcements/New-Feature.png) Our **[Global AI Operations Overview Dashboard](https://apipie.ai/dashboard){rel=""nofollow""}** just got a serious upgrade — and it's 🔥. We’ve already been grouping similar models into **expandable clusters** so you can easily compare them side-by-side, with the **top-performing model** in each group shown right on the dashboard. :br 👉 Be sure to **click expand** to dive deeper and see those comparisons in real-time. ## ::div{.text-center} ![Dashboard V2](https://apipie.ai/img/announcements/April-25/filters.png) :: ## 🔍 Now You Can Filter and Sort with Depth ### ✅ Filter Options: - **Model Types**: LLMs, voice, embeddings, and more. - **Subtypes**: multimodal, ChatX, pools, and other custom classifications. - **Latency Thresholds**: :br Exclude models with high latency — by prompt size or **TTFC** (Time To First Chunk/Token). - **Price Per Million Tokens**: :br Stay in budget *without* sacrificing performance. ### 📊 Sort Models By: - Throughput - Latency - Newest - Context size - Pricing :br ➡️ In ascending or descending order --- ## 📈 Need Availability Intel? We’ve got you covered with: - **30-day availability tracking** - **Visual latency graphs** And even more features are coming. 👀 --- ## 💸 Did You Know? Our **pricing reflects real-world costs** — calculated using our proprietary algorithm that evaluates the **actual cost of service per model, per provider**. :br*Nobody else reports pricing like this.* --- Stay tuned for more exciting features, fixes, and integrations! # Grok 3 & Grok 3 Mini: Models Now Available ![Grok 3](https://apipie.ai/img/announcements/New-Model.png) ## 🚀 **Grok 3 & Grok 3 Mini are now live on [APIpie.ai](https://apipie.ai){rel=""nofollow""}!** Just dropped: the latest flagship models from **xAI** — **Grok 3** and **Grok 3 Mini**, now fully integrated and ready to route via APIpie's unified AI API. 🔥 **Grok 3** > • Not a "thinking" model :br > • Optimized for structured tasks and benchmarks like **GPQA**, **LCB**, and **MMLU-Pro**:br > • Outperforms Grok 3 Mini on high-structure evals :br > • **Context window**: 131,072 tokens ⚡️ **Grok 3 Mini** > • High reasoning performance: **AIME'25: 83.0**, **AIME'24: 90.7**:br > • Transparent “thinking” traces included :br > • Defaults to low reasoning. > • **Context window**: 131,072 tokens 🛠️ **Available now** — plug in, route smart, and start building see it on our 👉 [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} # Internet Search Grounding: Web Data in Any AI Model ![Internet Grounding](https://apipie.ai/img/announcements/New-Feature.png) # 🚨 NEW FEATURE DROP — The Future of AI Has Arrived 🔥 Get ready to experience a whole new level of control and power — we’re unleashing **Search Grounding**, **Inline CLI**, **Multi-Tenancy**, and **real-time web-aware models** across our entire platform. If you’re building — or even just using — AI, **this changes everything**. ## ## 🌍 Search Grounding Is Now Live Now **ANY model** can answer with **real-time, verifiable information** — no plugins, agents, or browser tools needed. ### 🔎 Just drop in a live query, and we’ll: - Perform **high-quality searches** across the web - **Scrape and clean** top results - **Inject live info into the prompt** before the model call ✅ Works with **GPT-4o, Claude, LLaMA, Mistral** — literally **every LLM model we support**. --- ### ⚙️ Enable with a Simple Inline Config: ### 🛠️ Usage Options ```json "web_search_options": { "search_context_size": "medium" } ``` Use: - `"low"` – Light web grounding - `"medium"` – Moderate (\~3 results) - `"high"` – Deep (\~5 results) --- docs\Features\Internetsearch.md 📚 **Docs**: [Search + Scrape API Reference](https://apipie.ai/docs/features/internetsearch) # Inline CLI: Total API Control Inside the Prompt ::div{.text-center} ![Inline CLI Feature](https://apipie.ai/img/announcements/New-Feature.png) :: # 💻 Inline CLI — Total API Control Inside the Prompt Drop in **one-liner commands** to control everything from model selection to memory, integrity, shaping, and real-time search — **right from the user’s prompt**. ### 🧪 Examples: ```json :deepsearch :becreative :answerwithclaude :setmodel:openai/gpt-4o :setmemoryon ``` ### ⚙️ Enable via: ```json "inline_cli": "all" // (default) ``` ### 🗨️ User Prompt Example: ```json { "content": "Summarize today's AI news :deepsearch :setsearchlang:en :setsearchgeo:US" } ``` ### ✅ Supports: - **Model switching** & state-maintained overrides *(multi-tenancy-ready)* - **Search grounding** with region/language targeting - **Prompt shaping** (`:becreative`, `:beprecise`, etc.) - **Memory control** - **User state management** & more 💡 Just send `:help` as a user prompt to see inline help :br 🚫 Disable with: `"inline_cli": false` 📘 **Docs**: [Inline CLI Reference](https://apipie.ai/#) --- ### 🧠 Multi-Tenant Memory + Observability We now support **per-user state**, **usage tracking**, and **long-term memory** across all requests. 🔧 **Just include:** ```json "user": "your_user_id" ``` ### 🚀 This Unlocks: - 🔍 **Sub-user usage tracking** - 💾 **Memory tied to user + sub-user** - 🔐 **Persistent CLI settings** - 📊 **Deep observability by user/session** We’re launching one of the **most powerful multi-tenant AI backends** out there. --- ### 🧱 Works Seamlessly with OpenAI-Compatible APIs Everything above works **instantly** with your existing OpenAI-based stack using our `/v1/chat/completions` route. ✅ **Real-time web grounding**:br ✅ **Inline API controls**:br ✅ **Persistent memory**:br ✅ **CLI shaping & model overrides** --- 🧪 **Try it now**: [Docs](https://apipie.ai/docs/features/inlinecli) --- We’re not just leveling up — :br We’re launching a **whole new way to AI**. **Modular**, **memory-aware**, and **real-world grounded** — out of the box. :br Let’s build the future — smarter, faster, and together. — Team APIpie ⚙️ # OpenAI GPT-4.1: Increased Context & Performance ![ GPT-4.1](https://apipie.ai/img/announcements/New-Model.png) ## 🧠 **OpenAI Launches GPT-4.1** OpenAI has officially launched **GPT-4.1**, its latest flagship language model. This release includes three versions: :br**GPT-4.1**, **GPT-4.1 Mini**, and **GPT-4.1 Nano** — each with major improvements in coding, instruction following, and long-context reasoning. ⭐ **Key Features of GPT-4.1** - **Expanded Context Window**:br → Supports up to **1 million tokens**:br → Huge leap from GPT-4o's 128K tokens :br → Ideal for processing large datasets and complex reasoning tasks - **Improved Coding Capabilities**:br → **+21% over GPT-4o**, **+27% over GPT-4.5** in SWE-Bench :br → More accurate code generation and debugging - **Enhanced Instruction Following**:br → Better adherence to complex prompts :br → Reduced need for clarifications or re-prompts - **Cost Efficiency**:br → Operates at **26% lower cost** than GPT-4o :br → More affordable for developers and enterprise use - **Updated Knowledge Base**:br → Trained on data up to **June 2024**:br → Smarter, more relevant outputs 📡 **Availability & Transition** - **API Access Only**:br → Available via API for devs and businesses :br → No ChatGPT Plus access (yet) - **Model Phase-Outs**:br → GPT-4 will be removed from ChatGPT by **April 30**:br → GPT-4.5 preview deprecated by **July 14** 🛠️ **Available now** — plug in, route smart, and start building see it on our 👉 [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} # OpenAI o3 & o4-mini: Next-Gen Reasoning Models Now on APIpie ![ GPT-4.1](https://apipie.ai/img/announcements/New-Model.png) ## 🚀 **Now Live on APIpie.ai: OpenAI o3 + o4-mini** We're excited to roll out OpenAI’s most powerful reasoning models to date — available now through [APIpie.ai](https://apipie.ai){rel=""nofollow""}! ### 🔥 **OpenAI o3** > • 200K token context :br > • Full agentic tool use: Python, file analysis, browsing, image generation :br > • State-of-the-art reasoning across math, science, code, and visual tasks :br > • 20% fewer real-world errors than o1 ### ⚡ **OpenAI o4-mini** > • 200K token context :br > • Optimized for high-throughput, cost-efficient reasoning :br > • Best performance-to-price model on AIME 2024 & 2025 :br > • Tool-capable, with high-effort variants for complex logic ### 🧠 **What Makes These Models Special** > • Trained to reason about *when and how* to use tools :br > • Excels in code (SWE-bench, Codeforces), vision (CharXiv, MathVista), science (GPQA) :br > • Supports complex chains of reasoning + visual understanding :br > • Safer with reinforced refusal training and new frontier-risk evaluations ### 🧪 **Codex CLI** > Terminal-native LLM tool for advanced reasoning workflows. :br > Pass screenshots, low-res sketches, or code directly to the model. :br > Open-source now: [github.com/openai/codex](https://github.com/openai/codex){rel=""nofollow""} ::tip #### Use it with APIpie now by simply changing your ennvironment variables ```shell export OPENAI_BASE_URL="https://apipie.ai/v1" export OPENAI_API_KEY="your-APIpie-key-here" ``` :: ### 🔁 **Next Up** > • o3-pro coming soon :br > • Expanded Responses API support for tool calls + reasoning summaries :br > • Stay tuned for even smarter API routing & live performance benchmarking on APIpie! # Updates: Reporting, Real-Time Pricing & More ![Platform Improvements](https://apipie.ai/img/announcements/Announcement.png) ## 📢 Update Announcement 📢 We've rolled out several improvements and fixes to streamline functionality, enhance reporting, and improve the user experience. ## ✅ Model Reporting Updates - Image models and models with special characters now report latency properly - Coding models are now being actively tested for availability and latency metrics ## ⚙️ Real-Time Price Reporting We now provide real-time input/output cost reporting based on actual queries and real provider charges. This gives you the most precise breakdown of model usage costs. **Example (models route):** ```json { "enabled": 1, "available": 1, "type": "llm", "subtype": null, "provider": "bedrock", "id": "titan-text-premier-v1", "model": "titan-text-premier-v1", "route": "amazon.titan-text-premier-v1:0", "description": null, "max_tokens": 32000, "max_response_tokens": 32000, "latency": "7022/8810/395/na/na/5409", "query_count": 56, "img_price": null, "img_json": null, "avg_cost": 0.0002729346, "price_type": "token", "input_cost": 1.00123355, "output_cost": 1.49994697 } ``` **Details:** - **input\_cost** and **output\_cost** reflect real costs per million tokens (e.g., $1 per million input tokens, $1.50 per million output tokens) - If no specific cost data is available, the **avg\_cost** column provides an average cost per 1000 characters—useful but less precise ## 🌐 UI and SEO Improvements - Fixed the Robots file to improve web crawler behavior - Created a SiteMap for both the root site and documentation - Added Meta Tags to the root site - Implemented OG Schema & OG Image Generator to generate rich previews when sharing links - Included missing Meta Tags to ensure better indexing ## 🐞 Bug Fixes - Resolved an issue causing slow load times on the User Interface Thanks for your continued support! The APIpie Team 🥧 # Qwen-3: Advanced Multimodal AI Model ![Qwen-3](https://apipie.ai/img/announcements/New-Model.png) ## 🧠 **Alibaba Cloud Launches Qwen-3** Alibaba Cloud has released **Qwen-3**, an advanced multimodal AI model with enhanced reasoning capabilities, superior coding performance, and comprehensive vision understanding. The model represents a significant advancement in AI technology with improved instruction following and creative problem-solving abilities. 🚀 **Key Features of Qwen-3** - **Enhanced Reasoning**:br → Advanced logical thinking and problem-solving capabilities :br → Complex multi-step reasoning across various domains :br → Improved accuracy on mathematical and scientific tasks - **Superior Coding**:br → High-quality code generation in multiple programming languages :br → Advanced debugging and code analysis capabilities :br → Strong performance on programming benchmarks - **Multimodal Excellence**:br → Text, vision, and code understanding in one model :br → Advanced image analysis and visual reasoning :br → Comprehensive document and data processing - **Improved Instruction Following**:br → Precise adherence to complex user requirements :br → Better understanding of context and intent :br → Enhanced creative and analytical outputs - **Long Context Support**:br → Extended context windows for comprehensive understanding :br → Better handling of large documents and datasets :br → Improved coherence in long-form content generation 📡 **Performance & Capabilities** - **Coding Excellence**:br → Superior performance on HumanEval and MBPP benchmarks :br → Advanced code generation and optimization :br → Effective debugging and code analysis - **Reasoning Tasks**:br → Strong results on mathematical and logical problems :br → Advanced scientific reasoning capabilities :br → Excellent performance on complex problem-solving - **Multimodal Understanding**:br → Advanced vision-language comprehension :br → Document analysis and information extraction :br → Image understanding and description generation - **Multilingual Support**:br → Excellent performance across multiple languages :br → Strong cross-lingual understanding and translation :br → Cultural context awareness 🛠️ **Available now** — plug in, route smart, and start building see it on our 👉 [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} **Learn more:** [Qwen-3 Official Release](https://qwenlm.github.io/blog/qwen3/){rel=""nofollow""} # Phi-4: Microsoft's Small Language Models with Advanced Reasoning ![Phi-4](https://apipie.ai/img/announcements/New-Model.png) ## 🧠 **Microsoft Launches Phi-4 Reasoning Models** Microsoft has released **Phi-4-reasoning**, **Phi-4-reasoning-plus**, and **Phi-4-mini-reasoning** — marking a new era for small language models. These 14B and 3.8B parameter models achieve performance comparable to much larger models through advanced reinforcement learning and inference-time scaling. 🚀 **Key Features of Phi-4** - **Advanced Reasoning**:br → Inference-time scaling for complex multi-step problems :br → Chain-of-thought reasoning with detailed explanations :br → Trained on high-quality reasoning demonstrations - **Exceptional Efficiency**:br → **Phi-4-reasoning**: 14B parameters rivaling much larger models :br → **Phi-4-mini-reasoning**: 3.8B parameters for edge deployment :br → Lower computational costs and faster inference - **Superior Performance**:br → Better than OpenAI o1-mini and DeepSeek-R1-Distill-Llama-70B :br → **AIME 2025**: 78% (Phi-4-reasoning-plus) vs DeepSeek-R1's performance :br → Outperforms models 5x larger on reasoning benchmarks - **Open Weight Models**:br → **MIT license** for maximum deployment flexibility :br → Available on Azure AI Foundry and Hugging Face :br → No licensing fees for commercial use - **Multiple Variants**:br → **Phi-4-reasoning**: Base 14B reasoning model :br → **Phi-4-reasoning-plus**: Enhanced with RL for higher accuracy :br → **Phi-4-mini-reasoning**: Compact 3.8B model for resource-constrained environments 📡 **Performance & Benchmarks** - **Mathematical Excellence**:br → **AIME 2024**: 75.3% (reasoning), 81.3% (reasoning-plus) :br → **GPQA Diamond**: 65.8% (reasoning), 68.9% (reasoning-plus) :br → **OmniMath**: 76.6% (reasoning), 81.9% (reasoning-plus) - **Coding Capabilities**:br → **HumanEvalPlus**: 92.9% (reasoning), 92.3% (reasoning-plus) :br → Strong performance on programming tasks :br → Effective code generation and debugging - **General Reasoning**:br → Excellent instruction following and alignment :br → Strong performance across diverse problem types :br → Effective multi-step problem decomposition 🛠️ **Available now** — plug in, route smart, and start building see it on our 👉 [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} **Learn more:** [Microsoft Azure Phi-4 Blog](https://azure.microsoft.com/en-us/blog/one-year-of-phi-small-language-models-making-big-leaps-in-ai/){rel=""nofollow""} # Claude 4: Next-Gen AI Models with Advanced Coding ![Claude 4](https://apipie.ai/img/announcements/New-Model.png) ## 🧠 **Anthropic Launches Claude 4** Anthropic has released **Claude Opus 4** and **Claude Sonnet 4**, setting new standards for coding, advanced reasoning, and AI agents. Claude Opus 4 is the world's best coding model with sustained performance on complex tasks, while Claude Sonnet 4 delivers superior coding and reasoning with enhanced precision. 🚀 **Key Features of Claude 4** - **World's Best Coding Model (Opus 4)**:br → **72.5% on SWE-bench**, **43.2% on Terminal-bench**:br → Sustained performance on long-running tasks for hours :br → Thousands of steps with maintained focus and performance - **Enhanced Everyday Excellence (Sonnet 4)**:br → **72.7% on SWE-bench**, significant upgrade from Sonnet 3.7 :br → Balanced performance with optimal capability and efficiency :br → Enhanced steerability and instruction following - **Extended Thinking with Tool Use (Beta)**:br → Use web search and tools during reasoning :br → Alternate between thinking and tool use for better results :br → Multi-step workflows with real-time data access - **Revolutionary Capabilities**:br → **Parallel tool execution** for handling multiple tools simultaneously :br → **Enhanced memory** to extract and save key facts across sessions :br → **65% reduction** in shortcuts/loopholes compared to Sonnet 3.7 📡 **Performance & Industry Validation** - **Coding Excellence**:br → **SWE-bench Verified**: Opus 4 (72.5%), Sonnet 4 (72.7%) :br → **Terminal-bench**: Opus 4 (43.2%) - best performance :br → **TAU-bench**: Superior results on real-world agent scenarios - **Industry Leaders**:br → **Cursor**: State-of-the-art for coding with codebase understanding leap :br → **GitHub**: Powers new coding agent in GitHub Copilot :br → **Replit**: Improved precision for complex multi-file changes :br → **Rakuten**: 7-hour autonomous refactor with sustained performance - **Advanced Features**:br → **Memory capabilities**: Create and maintain memory files :br → **Tool-enhanced reasoning**: Better accuracy through tool use :br → **ASL-3 safety measures**: Higher AI Safety Level protocols 🛠️ **Available now** — plug in, route smart, and start building see it on our 👉 [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} **Learn more:** [Anthropic Claude 4 Announcement](https://www.anthropic.com/news/claude-4){rel=""nofollow""} # GPT-5 Features, API Changes, Integrations, and Benchmarks: The Complete Guide for Power Users and Developers ![GPT-5 Features and API Changes](https://apipie.ai/img/blog/August/GPT-5.svg) # GPT-5 Features, API Changes, Integrations, and Benchmarks: The Complete Guide for Power Users and Developers The long-awaited **GPT-5** has arrived, and it's much more than just a model upgrade. Whether you're a **power user** of ChatGPT or a **developer** integrating AI into complex workflows, GPT-5 delivers improvements in **accuracy, reasoning, integrations, and developer control** that change the game. This guide covers **GPT-5 features, API changes, benchmarks, Gmail and SharePoint integration details, verbosity vs. tokens, free-form tool calls**, and clears up the **context window size confusion** — optimized for both **AI enthusiasts** and **technical teams**. ## 🚀 What's New in GPT-5 OpenAI's GPT-5 introduces a **unified architecture** that blends fast responses, deep reasoning, and an intelligent router that selects the best strategy for each task. The result is **fewer errors, faster performance, and richer answers**. ### Key GPT-5 Features for All Users - **Enhanced reasoning** in math, coding, science, and multimodal tasks. - **65% fewer hallucinations** vs GPT-4 models. - **Safer completions** with more natural refusals. - **Transparent tool use** when integrations are active. ### Key GPT-5 API Features for Developers - **400K token context window** with **128K max output tokens** - massive capacity for complex tasks. - **Reasoning token support** for enhanced problem-solving capabilities. - **Predicted outputs** feature for faster response generation. - **Verbosity control** for guiding response richness without token micromanagement. - **Free-form tool calls** that return SQL, Python, CLI commands, or custom code formats instead of rigid JSON. - **Tool call preambles** that explain intent before execution. - **Better tool intent detection** for cleaner multi-step agent workflows. - **MCP (Model Context Protocol)** support for advanced integrations. ## 📊 GPT-5 Benchmarks vs GPT-4, GPT-4.1, and GPT-4.5 | Benchmark / Capability | GPT-4o | GPT-4.1 | GPT-4.5 | GPT-5 (Thinking Mode) | | ------------------------------- | -------- | ------- | ------- | --------------------- | | SWE-bench Verified (Coding) | \~54.6% | \~58% | — | **74.9%** | | Aider Polyglot (Multilang Code) | \~80% | \~82% | \~85% | **88%** | | GPQA (Science Reasoning) | \~70.1% | — | 71.4% | **85.7%** | | HealthBench Hard | \~30–35% | — | — | **46.2%** | | Hallucination Rate | \~11–15% | \~9% | \~8% | **4.8%** | **Takeaway:** GPT-5 consistently outperforms earlier models in reasoning-heavy, coding-intensive, and multilingual contexts. ## 📧 GPT-5 Gmail, SharePoint, and App Integrations **Power users** can now connect GPT-5 to Gmail, Google Calendar, Drive, and SharePoint directly inside ChatGPT's **Agent layer**. This enables: - Inbox summarization - Automated email drafting - Calendar event planning - Document searches **Developers**, however, don't get these connectors "for free" in the API. Instead: 1. Build your own integration with Gmail/SharePoint APIs. 2. Feed retrieved data to GPT-5. 3. Optionally define custom tools for automation. Result: **Instant setup** for end-users inside ChatGPT, **full customization** for developers via the GPT-5 API. ## 📏 Verbosity vs Max Tokens in GPT-5 Many confuse **verbosity** with `max_tokens` — but they're different: - **`max_tokens`**: A hard stop. If reached mid-sentence, the output ends abruptly. - **Verbosity**: A guidance signal. - **Low** = concise, high-signal answers. - **High** = in-depth explanations with examples. - Works with `max_tokens` but reduces the need for huge token limits just to ensure a detailed reply. For **power users**, this means more control over answer style. For **developers**, it simplifies prompt design for output length. ## 🔧 GPT-5 Free-Form Tool Calls Earlier OpenAI tool calls required **strict JSON outputs**, often leading developers to bypass them with prompt-engineered plain text. With GPT-5: - Tools can take **raw text arguments** (e.g., a SQL query or Python snippet). - Outputs are **self-contained** without conversational fluff. - **Tool call preambles** explain the action before execution. - Intent recognition is improved, so GPT-5 uses tools more reliably. Example: Instead of ```json { "query": "SELECT name, age FROM users WHERE active = TRUE;" } ``` you can receive directly: ```sql SELECT name, age FROM users WHERE active = TRUE; ``` This is a major win for developers who want clean, parse-ready outputs in agentic workflows. ## 🧠 GPT-5 Context Window: The Real Numbers The official specifications confirm GPT-5's impressive capacity: | Model | Context Window | Max Output Tokens | Pricing (per 1M tokens) | | ------ | --------------- | ----------------- | ------------------------ | | GPT-4o | 128K tokens | \~16K tokens | Lower cost | | GPT-5 | **400K tokens** | **128K tokens** | $1.25 input / $10 output | ### Key Specifications: - **400,000 context window** - 3x larger than GPT-4o - **128,000 max output tokens** - enables truly long-form generation - **October 1, 2024 knowledge cutoff** - most recent training data For **long-form generation** — books, full-codebase analysis, legal docs — GPT-5 removes the bottlenecks that plagued earlier models. ## 🎯 GPT-5 Key Takeaways for Power Users - More accurate answers with drastically fewer hallucinations. - Gmail, Calendar, Drive, and SharePoint integrations inside ChatGPT. - Verbosity gives better control over answer detail without abrupt cut-offs. ## 🔑 GPT-5 Key Takeaways for Developers - **400K token context** with **128K max output tokens** for massive document processing. - **Free-form tool calls** for natural, code-ready outputs. - **Reasoning token support** and **verbosity controls** for balancing speed and depth. - **Comprehensive tool ecosystem** with web search, code interpreter, and MCP support. - **Cost optimization** through tiered model options (GPT-5, GPT-5 mini, GPT-5 nano). ## 🚀 Get Started with GPT-5 via APIpie Ready to integrate GPT-5 into your workflows? APIpie provides seamless access to GPT-5 and other premium models: 1. **Easy Integration**: Connect to GPT-5 with a single API endpoint 2. **Cost Control**: Monitor usage with our comprehensive dashboard 3. **Developer Tools**: Access free-form tool calls and advanced features 4. **Multi-Model Support**: Switch between GPT-5, Claude, and other models Visit our [OpenAI Models documentation](https://apipie.ai/docs/models/openai) to get started with GPT-5 today. --- **Conclusion:** GPT-5 isn't just a step up from GPT-4.1 — it's a **smarter, more flexible, developer-friendly model**. For **power users**, it means deeper, faster, safer ChatGPT interactions. For **developers**, it's a toolkit for building next-gen agentic workflows, handling massive datasets, and delivering cleaner automation. --- *Stay updated with the latest AI model releases and integrations through APIpie. Check out our [Dashboard](https://apipie.ai/dashboard){rel=""nofollow""} for a real world view on AI.*