AIP overview

Institution: MIT

View original course

30 study materials · 8 sections

Palantir's Artificial Intelligence Platform (AIP) provides a secure and scalable framework for integrating Large Language Models (LLMs) with organizational data and operations. This course explores the core components of AIP, including AIP Logic, Agent Studio, and the Palantir Ontology, while emphasizing responsible AI practices and security. Students will learn to build, evaluate, and monitor AI-driven workflows, from prompt engineering to deploying custom models and agentic analysis tools.

Course Sections

AIP Foundations and Getting Started

Key concepts: AIP Logic · AIP Agent Studio · Ontology · Foundry and Apollo · Palantir Learning Portal

An introduction to the core architecture of AIP and the resources available to begin building AI-backed workflows.

AIP Foundations and Getting Started

Palantir's Artificial Intelligence Platform (AIP) represents a paradigm shift in enterprise software, moving beyond the "chatbot" model of AI to a deeply integrated orchestration layer. At its core, AIP is designed to bridge the gap between the non-deterministic reasoning of Large Language Models (LLMs) and the deterministic requirements of enterprise operations. It does not exist in a vacuum; rather, it is built upon the dual pillars of Foundry (the data plane) and Apollo (the deployment and orchestration plane), unified by the Palantir Ontology.

The Bedrock: Foundry and Apollo

To understand AIP, one must first master its architectural prerequisites. AIP is not a standalone product but an evolutionary layer on top of Palantir’s existing stack.

Foundry: The Data Plane

Foundry provides the "Data Foundation." It handles the ingestion, transformation, and governance of massive datasets. In the context of AIP, Foundry serves as the source of truth. Without a clean, governed data foundation, an LLM suffers from "garbage in, garbage out" (GIGO) at an enterprise scale.

Apollo: The Deployment Plane

Apollo is the continuous delivery and operations engine. It ensures that AIP can run across diverse environments—from public clouds (AWS, Azure, GCP) to on-premises data centers and even edge devices. Apollo manages the lifecycle of the models themselves, ensuring they are updated, secure, and performant.

Component Role in AIP Primary Function
Foundry Data Foundation Ingestion, cleaning, and semantic structuring of data.
Apollo Infrastructure Orchestration, versioning, and deployment across environments.
Ontology Semantic Layer Mapping raw data to real-world entities (Objects, Properties, Links).
AIP Intelligence Layer Reasoning, automation, and natural language interface.

The Semantic Brain: The Palantir Ontology

The Ontology is the most critical concept in the AIP ecosystem. While a standard database consists of tables and rows, the Ontology consists of Objects, Properties, and Links.

Definition: The Ontology The Ontology is a digital twin of the organization. It transforms kinetic data (raw logs, tables) into semantic entities (e.g., "Aircraft," "Maintenance Event," "Pilot") that an LLM can understand and manipulate.

Why the Ontology is Essential for AI

LLMs are historically poor at understanding relational database schemas. Asking an LLM to "Find all delayed flights caused by engine issues" requires it to know which tables to join, what the foreign keys are, and how "engine issues" are coded in a specific column.

In AIP, the LLM interacts with the Ontology. It sees an "Aircraft" object linked to a "Maintenance Log" object. This Grounding allows the LLM to reason about the business as a human would, significantly reducing hallucinations and increasing the precision of generated actions.

Mathematical Representation of Ontology Grounding

Let $O$ be the set of Objects, $P$ the set of Properties, and $L$ the set of Links. A business process is a subgraph $G' \subseteq G(O, L)$. When a user provides a natural language prompt $S$, the AIP system performs a mapping function:

f(S, G) \rightarrow A

where $A$ is a set of valid operations (Actions) defined within the Ontology's security constraints. This ensures that the AI cannot perform an operation that is not mathematically and logically defined within the system's bounds.


AIP Logic: Building Deterministic Workflows

AIP Logic is the development environment where engineers build LLM-backed functions. Unlike a standard script, an AIP Logic workflow is a mixture of structured logic and natural language reasoning.

How it Works: The Logic Pipeline

AIP Logic allows you to chain together multiple "blocks." These blocks can be:

  1. LLM Blocks: Where you provide a prompt and context to the model.
  2. Logic Blocks: Standard conditional statements or transformations.
  3. Ontology Blocks: Queries to the Object Set or executions of Actions.

Example: Automated Supply Chain Rerouting

Imagine a shipment is delayed. An AIP Logic function could:

  • Step 1: Identify the delayed Shipment object.
  • Step 2: Use an LLM block to summarize the reason for the delay from a raw text PDF (Bill of Lading).
  • Step 3: Query the Ontology for alternative Suppliers within a 500-mile radius.
  • Step 4: Use an LLM to draft a re-negotiation email based on the original contract terms.
# Conceptual representation of an AIP Logic Function
def handle_shipment_delay(shipment_id):
    # 1. Fetch Object from Ontology
    shipment = Ontology.get_object("Shipment", shipment_id)
    
    # 2. LLM Reasoning Block
    delay_reason = AIP.LLM.summarize(shipment.raw_logs, prompt="Why is this delayed?")
    
    # 3. Deterministic Query
    alternatives = Ontology.search("Suppliers").filter(lambda x: x.location == shipment.destination)
    
    # 4. Action Execution
    if "Weather" in delay_reason:
        AIP.execute_action("Update_Priority", shipment, priority="High")

AIP Agent Studio: Autonomous Operations

While AIP Logic is often used for specific, repeatable functions, AIP Agent Studio is designed for creating Agents—autonomous entities that can use "tools" to solve open-ended problems.

Agentic Workflows

An Agent in AIP is defined by its Tools (which are often AIP Logic functions or Ontology Actions) and its System Prompt. The Agent uses a "Reasoning Loop" (often following the ReAct framework: Reason + Act) to determine which tool to use next.

Feature AIP Logic AIP Agent Studio
Control Flow Explicit (Step-by-step) Autonomous (Goal-oriented)
Primary Use Case Predictable pipelines Ad-hoc problem solving
Flexibility Lower (Fixed paths) Higher (Dynamic paths)
Output Data or specific Action Multi-step resolution

The Model Context Protocol (MCP)

AIP utilizes the Palantir MCP to connect external AI systems and custom models directly to the Ontology. This model-agnostic approach allows organizations to swap the "brain" (e.g., moving from GPT-4 to Claude 3.5) without rebuilding the "nervous system" (the Ontology and Actions).

AI_DEMOI_DEMO--

AIP Evals: The Science of Non-Determinism

One of the greatest challenges in enterprise AI is Evaluation. Because LLMs are non-deterministic, a prompt that works today might fail tomorrow. AIP Evals provides a rigorous testing framework to quantify model performance.

The Evaluation Framework

AIP Evals introduces the concept of a Target Function and an Evaluation Function.

  • Target Function: The AIP Logic or Agent you are testing.
  • Evaluation Function: A separate, often more powerful LLM or a deterministic script that "grades" the output of the Target Function.

Metrics and P95 Duration

In AIP Evals, performance is measured across several dimensions:

  1. Correctness: Does the output match the "Golden Dataset" (ground truth)?
  2. Semantic Similarity: How close is the embedding vector of the output to the expected result?
  3. P95 Duration: The time it takes for 95% of requests to complete. This is critical for operational workflows where latency matters.
Metric Type Description
Exact Match Deterministic Boolean check against a known string/value.
LLM-as-a-Judge Probabilistic Using a model to score quality on a 1-5 scale.
Token Usage Cost Measuring the efficiency of the prompt.
Latency (ms) Performance Measuring the "wall clock" time of the execution.

Security, Governance, and Ethics

Palantir’s philosophy is that Responsible AI is not a feature—it is a foundational requirement. AIP integrates security directly into the data lineage.

Data Privacy and Third-Party LLMs

A common concern is that sensitive data will be used to train third-party models (like OpenAI or Google). AIP provides Technical and Contractual Guarantees:

  • No Retraining: Customer data is never used to train the base models of providers.
  • Regional Endpoints: Data can be restricted to specific geographic regions to comply with GDPR or CCPA.
  • Sensitive Data Scanner: An automated tool that intercepts prompts to check for PII (Personally Identifiable Information) before they leave the secure Palantir environment.

Explainability and Traceability

Every action taken by an AIP Agent is recorded in the Workflow Lineage. This allows for a full audit trail:

  • Which user triggered the agent?
  • What was the exact prompt sent to the LLM?
  • What was the "Chain of Thought" the LLM used?
  • Which Ontology Action was ultimately executed?

The Transparency Theorem For any AI-driven action $A$, there must exist a trace $T$ such that $T$ contains the set of all inputs $I$, the model version $M$, and the security context $C$ at time $t$. If $T$ is incomplete, the action $A$ is non-compliant.


Compute Usage and Model Management

Managing the cost of AI is a significant hurdle for many enterprises. AIP tracks usage through Compute-Seconds and Tokens.

Understanding Token Economics

LLMs process text in "tokens" (roughly 0.75 words). AIP usage is calculated based on:

  • Input Tokens: The context you provide (the prompt + Ontology data).
  • Output Tokens: The response generated by the model.

Because different models have different costs, AIP provides a Model Foundry where administrators can enable or disable specific models based on their cost-benefit profile. For example, an organization might use a small, cheap model for simple summarization but a large, expensive model for complex legal analysis.


Getting Started: The Palantir Learning Portal

To transition from theory to practice, Palantir provides the Learning Portal (learn.palantir.com). This is the primary resource for "AIP Speedruns"—intensive, hands-on tutorials designed to stand up a functional AI use case in hours.

Recommended Learning Path

  1. Foundry Foundations: Master the basics of data integration and the Ontology.
  2. AIP Logic Speedrun: Build your first LLM-backed function.
  3. Agent Configuration: Learn how to give an Agent tools and constraints.
  4. AIP Evals: Learn how to benchmark and "production-proof" your workflows.

Use Case Scoping

A critical part of getting started is Scoping. Not every problem is an AI problem. The Learning Portal provides frameworks to identify "High Value, High Feasibility" use cases, focusing on areas where the Ontology is already mature.

Security, Privacy, and AI Ethics

Key concepts: Zero Data Retention (ZDR) · Data Governance · Responsible AI · Explainability · Model Retraining Policy

Deep dive into the security protocols and ethical frameworks that govern AI usage within Palantir AIP.

Security, Privacy, and AI Ethics

In the enterprise landscape, the deployment of Large Language Models (LLMs) is often throttled not by a lack of capability, but by the gravity of risk. Palantir’s Artificial Intelligence Platform (AIP) addresses this by treating security, privacy, and ethics not as elective "add-ons," but as the foundational substrate upon which all AI operations are built. By integrating AI directly into the Ontology—the digital twin of an organization's data and logic—AIP ensures that every model interaction is governed by the same rigorous access controls and audit requirements as the underlying data itself.

AI_SVGI_SVGThe AIP Security & Ethics Architecture: Illustrating the flow from raw data through the Ontology, the Governance Proxy layer, and finally to the LLM providers, highlighting the Zero Data Retention (ZDR) boundary.*

The Foundation: Zero Data Retention (ZDR) and Retraining Policies

The primary inhibitor for enterprise AI adoption is the "data leakage" problem: the fear that proprietary data sent to a model provider will be used to train future iterations of that model, eventually surfacing in a competitor’s query results. AIP mitigates this through a combination of technical architecture and strict contractual enforcement.

Zero Data Retention (ZDR)

Definition: Zero Data Retention (ZDR) is a security guarantee where the LLM provider (e.g., OpenAI, Anthropic, Google) agrees to never store the prompt or completion data on persistent disk beyond the immediate duration of the inference request, and never utilizes that data for model training or improvement.

In AIP, ZDR is the default posture for integrated third-party models. When a user interacts with an LLM via AIP Logic or Agent Studio, the data is encrypted in transit using industry-standard protocols (TLS 1.2+). Once the request reaches the provider's specialized "ZDR endpoint," the data remains in volatile memory only for the time required to generate a response.

Model Retraining Policy

Palantir maintains a strict Model Retraining Policy that serves as a legal and technical firewall. While consumer-grade AI tools often use "opt-out" mechanisms for data training, AIP's enterprise-grade integrations are "opt-in" by default only for the customer's own private fine-tuning, should they choose to pursue it. For standard operations, the policy ensures:

  1. No Cross-Pollination: Data from Enrollment A never influences the weights of a model used by Enrollment B.
  2. No Provider Learning: Third-party providers are contractually barred from using any Palantir-routed data to improve their base models.
Feature Standard Consumer API AIP Enterprise Proxy (ZDR)
Data Storage Often 30+ days for "abuse monitoring" Zero persistent storage
Model Training Data may be used for RLHF/Training Explicitly prohibited
Encryption Standard TLS TLS + Managed Identity + Bearer Tokens
Audit Log Limited to provider logs Full distributed tracing in Foundry
Regionality Often global/US-centric Configurable Regional Endpoints

Data Governance and The Proxy Architecture

AIP does not allow LLMs to "crawl" raw data. Instead, it utilizes a Proxy Architecture that intercepts all calls to LLM providers. This proxy serves as the enforcement point for Data Governance.

The Role of the Ontology

The Ontology acts as the semantic translator between the LLM and the organization's data. By forcing AI agents to interact with the Ontology rather than raw tables, AIP ensures:

  • Object-Level Security: If a user does not have permission to see "Salary Data," the AI agent acting on their behalf cannot access it either.
  • Purpose-Based Access: Data is only surfaced to the model if it is relevant to the specific task, minimizing the "surface area" of data exposure.

LLM-Provider Compatible APIs

AIP provides proxy endpoints that are compatible with native SDKs (like the OpenAI Python library). This allows developers to use familiar tools while benefiting from Foundry’s security stack.

# Example of a secure, governed LLM call via the AIP Proxy
from palantir_models.sdk import ModelProvider

# The client is automatically configured with Foundry's 
# Bearer Token and Data Governance headers.
client = ModelProvider.get_client("openai-gpt-4o")

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Analyze the Q3 churn logic in the Ontology."}],
    # AIP intercepts this call to ensure:
    # 1. The user has 'read' access to the churn objects.
    # 2. The request is routed to a ZDR-compliant regional endpoint.
    # 3. The execution is logged for auditability.
)

print(response.choices[0].message.content)

Responsible AI: The Five Pillars of Ethical Deployment

Palantir’s framework for Responsible AI (RAI) moves beyond theoretical ethics into operationalized software features. This framework is categorized into five core themes.

1. Equity (Fairness)

AI models can inadvertently reflect biases present in their training data. AIP addresses this through Sub-population Analysis within AIP Evals. Users can run test cases across different demographic or geographic segments to ensure the model's performance is consistent and unbiased.

2. Explainability (Transparency)

Definition: Explainability in AI refers to the ability to trace the reasoning path of a model, identifying which specific data points or logic steps led to a particular output.

AIP provides transparency through AIP Logic and Distributed Tracing. When an agent performs a task, AIP records the "Chain of Thought," showing exactly which Ontology objects were queried and which tools were invoked. This transforms the "black box" of an LLM into a "glass box" where every decision is auditable.

3. Reliability (Safety)

To ensure reliability, AIP utilizes AIP Evals, a dedicated testing suite. Because LLMs are non-deterministic (they can give different answers to the same prompt), reliability is measured through:

  • P95 Duration: Monitoring the latency of AI responses to ensure operational stability.
  • Deterministic Guardrails: Using "AIP Logic" to wrap LLM calls in hard-coded logic gates, ensuring the model cannot deviate from prescribed business rules.

4. Traceability (Auditability)

Every interaction in AIP is logged. This includes the prompt, the completion, the user ID, the timestamp, and the specific version of the model used. This Execution History is typically maintained for 30 days in active logs and can be archived for long-term compliance.

5. Collaborative (Interdisciplinary)

AIP is designed to be used by "Human-in-the-Loop" systems. It does not replace the human; it augments them. Features like AIP Threads allow humans to review, comment on, and override AI-generated suggestions, ensuring that accountability always rests with a person.

Pillar Technical Implementation in AIP Goal
Equity AIP Evals + Subset Testing Minimize algorithmic bias
Explainability Traceability + Ontology Metadata Understand the "Why"
Reliability AIP Evals + P95 Monitoring Ensure consistent performance
Traceability Apollo-managed logs + Execution History Regulatory compliance and auditing
Accountability Human-in-the-loop (AIP Threads) Maintain human oversight

AIP Evals: Operationalizing Ethics and Performance

The non-deterministic nature of LLMs—where the same input can yield different outputs—poses a significant challenge for security and ethics. AIP Evals is the solution to this "hallucination" and "drift" problem.

How AIP Evals Works

  1. Test Cases: A library of "Golden Inputs" (known queries) and "Expected Outputs."
  2. Target Function: The specific AIP Logic or Agent being tested.
  3. Evaluation Function: A secondary model or a set of code-based rules that "grades" the output of the target function.
  4. Metrics: Quantitative scores (0-1) for accuracy, tone, safety, and adherence to constraints.

Concrete Example: Detecting Sensitive Data

A common ethical pitfall is the accidental inclusion of PII (Personally Identifiable Information) in model prompts. AIP includes a Sensitive Data Scanner that can be integrated into the evaluation pipeline. If a prompt contains a pattern matching a Social Security Number or a private API key, the scanner flags it, and the evaluation fails, preventing that logic from being deployed to production.

AI_DEMOI_DEMOInteractive Simulation: A user adjusts a "Temperature" slider and observes how the "Reliability Score" in AIP Evals fluctuates. As temperature increases (more creativity), the adherence to the "Ontology Schema" decreases, demonstrating the trade-off between LLM flexibility and deterministic safety.*

Observability and Performance Monitoring

Security and ethics also encompass the "Health" of the system. AIP Observability provides deep insights into the lifecycle of an AI request.

  • Metrics and P95: High latency (P95 duration) can indicate a model is struggling with complex reasoning or that a provider's endpoint is degraded.
  • Distributed Tracing: If an AIP Agent fails, distributed tracing allows an engineer to see if the failure happened at the LLM provider, the Ontology query layer, or the data transformation step.
  • Compute Usage: AIP tracks usage in Tokens and Compute-seconds. This is critical for preventing "Denial of Wallet" attacks, where inefficient prompts or infinite loops drain an organization’s compute budget.
Metric Description Ethical/Security Significance
Token Count Number of units processed Cost control and resource equity
P95 Duration 95th percentile of response time System reliability and user trust
Error Rate % of failed model calls Identifying model drift or provider instability
Ontology Hit Rate % of prompts successfully grounded in data Measuring "Hallucination" levels

Common Pitfalls and Misconceptions

Despite the robust framework of AIP, users must be aware of subtle edge cases:

  1. The "Prompt Injection" Fallacy: Some believe that ZDR prevents prompt injection. It does not. ZDR prevents data retention, but the model can still be "tricked" into ignoring instructions during a single session. This is why AIP Logic (which uses sequential, constrained steps) is safer than a single open-ended chat window.
  2. Georestrictions: Even with ZDR, data residency laws (like GDPR) may require that data never leaves a specific region. Users must ensure they are using Regional Endpoints (e.g., AWS Frankfurt or Azure Netherlands) rather than the default global endpoints.
  3. Model Agnosticism vs. Model Consistency: AIP is model-agnostic, meaning you can switch from GPT-4 to Claude 3. However, a prompt that is "safe" and "ethical" for one model might produce "hallucinations" in another. Continuous re-evaluation in AIP Evals is mandatory when switching models.

Summary of Governance Workflow

The lifecycle of a secure AI application in AIP follows a rigorous path:

  1. Scoping: Define the use case and identify sensitive data via the Sensitive Data Scanner.
  2. Ontology Integration: Map the data to the Ontology to inherit existing access controls.
  3. Development: Build the logic in AIP Logic or Agent Studio, utilizing prompt engineering best practices (Clarity, Specificity, Constraints).
  4. Evaluation: Run the logic through AIP Evals to check for bias, accuracy, and reliability.
  5. Deployment: Enable the model via the Control Panel, ensuring ZDR and regionality are enforced.
  6. Monitoring: Use AIP Observability to track performance and audit every interaction.

"Responsible AI is not an afterthought; it is fundamental to how we build technology. Our approach centers on developing software that enables responsible AI use throughout the entire system lifecycle." — Palantir AI Ethics & Governance Philosophy.

AI_FLASHCARDSI_FLASHCARDS Zero Data Retention (ZDR): A guarantee that LLM providers do not store or train on your data.

  • Ontology: The semantic layer that enforces data permissions for AI agents.
  • AIP Evals: The testing suite used to measure model accuracy and ethical alignment.
  • P95 Duration: A metric used to track the reliability and performance of AI workflows.
  • Proxy Endpoint: The secure gateway that intercepts and governs all calls to third-party LLMs.
  • Execution History: The 30-day audit trail of all AI interactions and logic steps.

AI_QUIZI_QUIZ. Which feature in AIP prevents an LLM from using your proprietary data to train its next model?

  • A) AIP Logic
  • B) Zero Data Retention (ZDR)
  • C) Distributed Tracing
  • D) Vega Visualizations (Correct: B)
  1. How does the Ontology improve AI security?

    • A) By encrypting the model weights.
    • B) By providing a semantic layer that respects existing object-level access controls.
    • C) By increasing the speed of token generation.
    • D) By replacing the need for an LLM provider. (Correct: B)
  2. What is the purpose of AIP Evals?

    • A) To visualize data in 3D.
    • B) To manage the non-deterministic nature of LLMs through benchmarking and testing.
    • C) To store user passwords.
    • D) To translate natural language into SQL. (Correct: B)
  3. True or False: AIP allows you to use open-source SDKs while still enforcing Foundry's data governance.

    • (Correct: True, via LLM-provider compatible proxy endpoints)

AI_STUDY_GUIDEI_STUDY_GUIDE*Key Concepts to Master:**

  • ZDR Mechanics: Understand the difference between data in transit, data in volatile memory, and persistent storage.
  • The RAI Framework: Be able to list and explain the five pillars (Equity, Explainability, Reliability, Traceability, Collaborative).
  • AIP Evals Workflow: Know how to set up a test case, a target function, and an evaluation function to detect model drift.
  • Governance Proxy: Understand how Bearer Tokens and regional endpoints ensure data residency and auditability.
  • Ontology-Grounded Prompting: Explain why querying the Ontology is more secure than providing raw data to an LLM.
Security, Privacy, and AI Ethics - AIP overview - diagram 1
Security, Privacy, and AI Ethics - AIP overview - diagram 1

LLM Infrastructure and Capacity Management

Key concepts: Model-Agnosticism · Compute-seconds · Tokens Per Minute (TPM) · Requests Per Minute (RPM) · Proxy Endpoints

Technical details on how AIP manages various LLM providers, compute costs, and rate limits.

LLM Infrastructure and Capacity Management

In the contemporary enterprise landscape, the transition from experimental "sandbox" AI to production-grade operational intelligence requires more than just a high-performing model. It necessitates a robust, scalable, and governed infrastructure capable of mediating between heterogeneous Large Language Model (LLM) providers and the complex data environments of the modern organization. Palantir’s Artificial Intelligence Platform (AIP) addresses this by providing a sophisticated orchestration layer that abstracts the underlying hardware and model-specific complexities, allowing developers to focus on logic and workflow integration.

This article provides a deep-dive into the mechanics of LLM infrastructure within AIP, focusing on the mathematical underpinnings of capacity management, the architectural significance of model-agnosticism, and the governance frameworks that ensure security and cost-efficiency.

AI_SVG--

Model-Agnosticism: The Abstraction of Intelligence

Model-Agnosticism is the architectural principle that decouples the application layer (logic, prompts, and workflows) from the specific LLM implementation (e.g., GPT-4, Claude 3.5, Gemini 1.5). In the context of AIP, this means that a single workflow—such as an AIP Logic function or an Agentic workflow—can be retargeted to a different model with minimal to no code changes.

Why It Matters

The AI landscape is characterized by rapid volatility. A model that is the "state-of-the-art" (SOTA) today may be superseded by a more efficient or capable competitor within months. Without a model-agnostic framework, organizations face "vendor lock-in," where their entire intellectual property (IP) regarding prompts and data integration is tied to a specific provider’s API and quirks. Agnosticism provides:

  1. Future-Proofing: Easy migration to newer, better models.
  2. Cost Optimization: Routing simple tasks to cheaper models and complex reasoning to "frontier" models.
  3. Redundancy: Switching providers in the event of regional outages or API rate-limiting.

The Model Foundry and Enablement

AIP manages this through the Model Foundry, a centralized hub where administrators enable specific models for use across the enrollment. Each model undergoes a state-based lifecycle:

State Description Usage Impact
Available The model is supported by the platform but not yet enabled for the specific enrollment. Cannot be used in workflows.
Enabled The model is active and accessible to developers. Consumes compute-seconds; available in AIP Logic/Studio.
Deprecated The model is being phased out by the provider. Existing workflows continue; new ones should avoid it.
Retired The model is no longer functional. Workflows using this model will fail until updated.

Capacity Management: TPM and RPM

Managing the throughput of an LLM is fundamentally different from managing traditional compute (CPU/RAM). Because LLMs are hosted as managed services (often by third parties like OpenAI or Anthropic), capacity is governed by Rate Limits. These limits are expressed through two primary metrics: Tokens Per Minute (TPM) and Requests Per Minute (RPM).

Tokens Per Minute (TPM)

TPM is a measure of the raw volume of data processed by the model. A "token" is the fundamental unit of text for an LLM, roughly equivalent to 0.75 words.

Definition: Let $T_{in}$ be the number of input tokens and $T_{out}$ be the number of output tokens. For a given window of 60 seconds, the total consumption $C_{T}$ must satisfy: $$\sum_{i=1}^{n} (T_{in, i} + T_{out, i}) \leq TPM_{limit}$$ where $n$ is the number of requests in that minute.

Requests Per Minute (RPM)

RPM measures the concurrency or frequency of calls. Even if your requests are very short (low tokens), sending too many in quick succession will trigger rate limiting.

Definition: Let $R$ be the count of individual API calls. The constraint is simply: $$R_{60s} \leq RPM_{limit}$$

The Relationship Between TPM and RPM

In production, these two limits act as a "dual-throttle."

  • High TPM / Low RPM: Suitable for batch processing large documents (e.g., summarizing a 50-page PDF).
  • Low TPM / High RPM: Suitable for interactive chat applications with many users sending short messages.
Metric Primary Constraint Bottleneck Scenario
TPM Model context window and processing bandwidth. Large-scale data extraction from long documents.
RPM API Gateway overhead and connection handling. High-concurrency "micro-tasks" (e.g., classifying 1,000 rows).

Common Pitfall: The "Burstiness" Problem

A common mistake is assuming that a TPM of 60,000 allows for a single request of 60,000 tokens at the start of the minute. Most providers implement a "leaky bucket" algorithm where capacity is replenished continuously. Sending a massive burst can trigger a 429 Too Many Requests error even if the total minute-long quota hasn't been met.


Compute-seconds: The Economic Unit of LLMs

To provide a unified billing and resource tracking mechanism across different providers, AIP uses the concept of Compute-seconds. Since different models have vastly different costs (e.g., GPT-4o is significantly more expensive than GPT-3.5 Turbo), a raw token count is insufficient for financial governance.

The Derivation of Compute-seconds

AIP normalizes model usage by assigning a "rate" to tokens based on the model's complexity and the provider's pricing.

Formula: The total compute usage $U$ for a single request is calculated as: $$U = (T_{in} \times R_{in}) + (T_{out} \times R_{out})$$ Where:

  • $T_{in}, T_{out}$ are the counts of input and output tokens.
  • $R_{in}, R_{out}$ are the model-specific rates (expressed in compute-seconds per million tokens).

Worked Example: Cost Calculation

Consider two models: Model Alpha (High Reasoning) and Model Beta (Fast/Cheap).

Model Input Rate ($R_{in}$) Output Rate ($R_{out}$)
Alpha 10.0 30.0
Beta 0.5 1.5

Scenario: A user processes a prompt of 1,000 tokens and receives a 500-token response.

  1. Using Model Alpha: $U = (1000 \times 10.0 / 10^6) + (500 \times 30.0 / 10^6)$ $U = 0.01 + 0.015 = 0.025$ compute-seconds.
  2. Using Model Beta: $U = (1000 \times 0.5 / 10^6) + (500 \times 1.5 / 10^6)$ $U = 0.0005 + 0.00075 = 0.00125$ compute-seconds.

Insight: Model Alpha is 20x more expensive for the same task. Infrastructure management in AIP allows administrators to set quotas on these compute-seconds to prevent runaway costs during development.

AI_DEMOI_DEMO--

Proxy Endpoints: Governed Connectivity

For developers accustomed to using native SDKs (like the openai Python library or langchain), AIP provides Proxy Endpoints. These are Foundry-hosted URLs that mimic the API structure of the underlying provider while routing the traffic through Palantir’s security and governance stack.

How Proxy Endpoints Work

  1. Authentication: The developer uses a Foundry API token instead of a third-party API key.
  2. Interception: The request hits the Foundry Proxy.
  3. Governance Check: The proxy verifies if the user has permission to use the model and if the enrollment has remaining TPM/RPM capacity.
  4. Audit Logging: The request and response metadata (not the content, depending on privacy settings) are logged for observability.
  5. Forwarding: The proxy strips the Foundry headers, adds the provider’s credentials, and forwards the request to the actual LLM endpoint (e.g., Azure OpenAI or AWS Bedrock).

Implementation Example

Using a standard OpenAI-compatible client to connect to AIP:

from openai import OpenAI

# The client points to the Foundry Proxy instead of direct OpenAI
client = OpenAI(
    base_url="https://<foundry-url>/api/v1/proxy/foundry-ml/openai",
    api_key="your-foundry-token"
)

response = client.chat.completions.create(
    model="gpt-4o", # This refers to the model enabled in Model Foundry
    messages=[{"role": "user", "content": "Analyze the quarterly logistics data."}]
)

print(response.choices[0].message.content)

Why Use Proxies?

  • Data Privacy: The proxy ensures that data sent to third-party providers is governed by the organization's legal agreements (e.g., ensuring data is not used for model retraining).
  • Regionality: Proxies can enforce georestrictions, ensuring that a user in the EU only hits model endpoints located within the EU to comply with GDPR.
  • Unified Observability: All calls, regardless of the library used, show up in the AIP Observability dashboard.

Observability and Performance Monitoring

In an LLM-backed system, performance is not just about "up or down." It is about latency distributions and failure modes. AIP Observability provides a suite of tools to monitor the health of the LLM infrastructure.

Key Metrics

  • P95 Duration: The time it takes for the slowest 5% of requests to complete. LLMs are notoriously variable in latency; monitoring the P95 is crucial for maintaining a good user experience.
  • Token Throughput: Tracking TPM usage over time to identify peak hours and potential capacity exhaustion.
  • Error Rates: Tracking 429 (Rate Limited), 401 (Unauthorized), and 5xx (Provider Down) errors.

Distributed Tracing

Because AIP workflows often involve multiple steps (e.g., an Agent searching the Ontology, then calling an LLM, then writing back to the Ontology), AIP uses Distributed Tracing. This allows a developer to see the entire "lineage" of a request. If an Agent is slow, tracing reveals whether the bottleneck was the LLM generation, the data retrieval from the Ontology, or a complex SQL transformation.


Enrollment Tiers and Scaling

To accommodate different organizational needs, AIP infrastructure is often organized into Enrollment Tiers. These tiers define the baseline capacity (TPM/RPM) and the level of support for high-availability workloads.

Tier Typical TPM Use Case Scaling Mechanism
Medium 50k - 100k Prototyping and small team tools. Shared multi-tenant pools.
Large 250k - 500k Department-wide operational apps. Dedicated capacity allocations.
XL / Enterprise 1M+ Mission-critical, automated pipelines. Multi-region, provisioned throughput.

Scaling Logic

When an enrollment reaches its limit, AIP can implement Priority Queuing. Critical production workflows (e.g., an automated supply chain alert) can be given priority over ad-hoc analyst queries, ensuring that infrastructure constraints do not break operational continuity.


Common Pitfalls in Capacity Management

  1. Ignoring Output Token Costs: Developers often optimize the prompt (input) but forget that the LLM's response (output) is often 3x-5x more expensive in terms of compute-seconds.
  2. Hard-coding Model IDs: Using specific model strings (e.g., gpt-4-0613) in code rather than using AIP's model aliases makes migrations difficult when that specific version is retired.
  3. Lack of Retry Logic: Even with high TPM, transient network errors or provider-side spikes occur. Robust infrastructure requires exponential backoff and retry logic, which is partially handled by AIP but must be considered in custom code.
  4. Over-reliance on "Frontier" Models: Using a high-reasoning model for a simple classification task wastes compute-seconds and increases latency.

AI_QUIZI_QUIZ--

AI_STUDY_GUIDEI_STUDY_GUIDE### Key Terms Summary

  • Model-Agnosticism: The ability to switch LLM providers without rewriting application logic.
  • Compute-seconds: The normalized unit of cost in AIP, calculated from input/output tokens and model rates.
  • TPM/RPM: The two primary throttles for LLM capacity (Tokens/Requests Per Minute).
  • Proxy Endpoint: A governed gateway that allows standard AI tools to interact with Foundry-managed models.
  • Model Foundry: The administrative interface for enabling, deprecating, and managing LLM access.

Core Formulas to Remember

  1. Capacity Constraint: $\text{Tokens}{in} + \text{Tokens}{out} \leq \text{TPM Limit}$
  2. Usage Cost: $U = \sum (\text{Tokens} \times \text{Rate})$
  3. Token Conversion: $1 \text{ Token} \approx 0.75 \text{ Words}$

Implementation Checklist

  • Are the required models enabled in the Model Foundry?
  • Does the enrollment have sufficient TPM for the expected user load?
  • Are Proxy Endpoints configured for developers using external SDKs?
  • Is AIP Observability being used to monitor P95 latency?
  • Have quotas been set for compute-second consumption to manage budget?

Best Practices for Prompt Engineering

Key concepts: Clarity and Specificity · Few-shot Prompting · Iterative Refinement · Constraint Setting

Strategies for designing effective inputs to optimize LLM performance and reliability.

Best Practices for Prompt Engineering

Overview

Prompt engineering is the systematic and iterative process of designing, refining, and optimizing inputs—known as prompts—to guide Large Language Models (LLMs) toward generating high-quality, accurate, and contextually relevant outputs. Within the context of Palantir’s Artificial Intelligence Platform (AIP), prompt engineering transcends simple "chatbot" interactions; it becomes a core engineering discipline used to build robust, deterministic-like workflows out of non-deterministic models.

In the AIP ecosystem, prompt engineering is the bridge between the Ontology (your organization's digital twin) and AIP Logic or Agent Studio. By mastering the nuances of how LLMs interpret natural language, engineers can automate complex operational processes, transform raw data into structured insights, and ensure that AI-driven actions remain within the bounds of corporate governance and security.

AI_SVGI_SVG--

Clarity and Specificity

What it is

Clarity and Specificity refers to the elimination of ambiguity in a prompt by providing explicit instructions, defining the model's persona, and detailing the exact context of the task. Mathematically, this can be viewed as narrowing the probability distribution of the model's output tokens toward a specific, desired subset of the latent space.

Why it matters

LLMs are trained on vast, heterogeneous datasets. Without specific constraints, a model may default to a "generalist" tone or interpret a command in multiple ways. In an enterprise environment, ambiguity leads to "hallucinations" or irrelevant responses that can break downstream data pipelines. Specificity ensures that the model's "attention" is focused on the relevant variables.

How it works: The Anatomy of a Clear Prompt

A high-signal prompt typically contains four components:

  1. Role/Persona: Assigning a professional identity (e.g., "You are a Senior Supply Chain Analyst").
  2. Task: A clear, verb-centric instruction (e.g., "Extract," "Summarize," "Validate").
  3. Context: The specific data or background information (often pulled from the Palantir Ontology).
  4. Output Format: The desired structure (e.g., "Return a JSON object with keys 'id' and 'status'").
Component Purpose Example
Persona Sets the tone and domain expertise "Act as a Lead Maintenance Engineer."
Instruction Defines the primary action "Analyze the following sensor logs for anomalies."
Context Provides the 'ground truth' "The logs cover the last 24 hours for Turbine-X."
Format Ensures machine-readability "Output the results as a Markdown table."

Concrete Example

Weak Prompt: "Tell me about the late shipments." Strong Prompt: "You are a Logistics Coordinator. Review the attached list of 'Shipment' objects from the Ontology. Identify all shipments where the actual_delivery_date is later than the promised_delivery_date. For each late shipment, provide the tracking_id and a 10-word summary of the delay reason found in the comments field. Format the output as a CSV."

Common Pitfalls

  • Over-reliance on adjectives: Using words like "better" or "faster" is subjective. Use quantitative benchmarks instead.
  • Instruction Bloat: Providing too many conflicting instructions in a single paragraph, which dilutes the model's focus.

Few-shot Prompting

What it is

Few-shot Prompting is a technique where the model is provided with a small number of high-quality examples (shots) within the prompt to demonstrate the desired input-output mapping. This leverages the model's In-Context Learning (ICL) capabilities without requiring any weight updates or fine-tuning.

Why it matters

Zero-shot prompting (asking without examples) often fails when the task requires a highly specific format, a particular "voice," or complex logic that is difficult to describe in prose alone. Few-shot prompting provides a "pattern" for the model to follow, significantly increasing the reliability of structured outputs like SQL queries, JSON, or domain-specific languages (DSL).

How it works: The Pattern Match

The model uses the provided examples to infer the underlying transformation logic. The effectiveness of few-shot prompting is highly dependent on the diversity and relevance of the examples.

Definition: In-Context Learning (ICL) The ability of a pre-trained LLM to learn a new task at inference time by simply observing examples in its context window, without any gradient-based updates to its parameters.

Shot Count Type Use Case
0-shot Zero-shot Simple, common tasks (e.g., "Translate 'Hello' to French").
1-shot One-shot Tasks with a clear but unique format requirement.
3-5 shots Few-shot Complex transformations, sentiment analysis, or code generation.

Concrete Example in AIP Logic

In AIP Logic, you might use few-shot prompting to help a model map natural language to an Ontology action.

### Task: Map user requests to the 'Update Inventory' action.

### Examples:
Request: "We just received 50 units of Grade A steel for Warehouse 4."
Action: {"action": "update_inventory", "params": {"quantity": 50, "material": "Steel_A", "loc": "WH4"}}

Request: "Subtract 10 broken valves from the stock in the North Wing."
Action: {"action": "update_inventory", "params": {"quantity": -10, "material": "Valve", "loc": "North_Wing"}}

### Current Request:
Request: "Add 100 boxes of filters to the main depot."
Action:

Variations / Extensions

  • Dynamic Few-shot: Using a vector database to retrieve the most relevant examples for a specific query and injecting them into the prompt at runtime.
  • Chain-of-Thought (CoT) Few-shot: Providing examples that include the "reasoning steps" before the final answer.

Iterative Refinement and AIP Evals

What it is

Iterative Refinement is the "scientific method" applied to prompt engineering. It involves a continuous loop of:

  1. Drafting a prompt.
  2. Executing it against a representative dataset.
  3. Evaluating the output against success metrics.
  4. Adjusting the prompt based on observed failures.

Why it matters

LLMs are non-deterministic; the same prompt can yield different results across different runs or different models (e.g., switching from GPT-4o to Claude 3.5). Within Palantir AIP, AIP Evals provides the infrastructure to quantify this performance, moving prompt engineering from "vibes-based" development to data-driven engineering.

How it works: The Evaluation Suite

AIP Evals allows users to define a Target Function (the prompt/logic being tested) and an Evaluation Function (the grader).

  1. Test Cases: A collection of inputs and expected "ground truth" outputs.
  2. Metrics: Quantitative measures like Accuracy, F1 Score, or "LLM-as-a-judge" scores.
  3. Comparison: Running the same test cases across different versions of a prompt or different models to see which performs better.
Metric Type Description Best For
Exact Match Does the output match the ground truth string exactly? Code, IDs, Boolean values.
Semantic Similarity Does the output mean the same thing as the ground truth? Summarization, Q&A.
LLM Grader A second LLM evaluates the output based on a rubric. Tone, safety, reasoning quality.
P95 Duration The time it takes for 95% of requests to complete. Performance and latency monitoring.

AI_DEMOI_DEMO### Common Pitfalls

  • Overfitting to a single example: Changing a prompt to fix one error, only to break five other cases. This is why a diverse Evaluation Suite is critical.
  • Ignoring Latency: A prompt that is 2,000 words long might be 1% more accurate but 5x slower and more expensive.

Complexity Management and Sequential Prompting

What it is

Complexity Management involves breaking down a multifaceted problem into smaller, sequential sub-tasks. Instead of asking an LLM to "Analyze this 50-page document and write a 10-page risk report," you break the task into discrete steps: extract key facts, identify risks per category, and then synthesize the report.

Why it matters

LLMs have a limited "reasoning capacity" per forward pass. As the complexity of the prompt increases, the likelihood of the model skipping steps or losing track of constraints (the "lost in the middle" phenomenon) increases. Sequential prompting—often implemented via AIP Logic or Agentic Workflows—allows each step to have its own specialized prompt and context.

How it works: Chain of Thought (CoT)

One of the most powerful ways to manage complexity is to force the model to "think" before it acts. By adding the instruction "Let's think step-by-step," you encourage the model to generate intermediate reasoning tokens, which improves performance on logical and mathematical tasks.

Key Insight: The Token-Reasoning Correlation There is a direct correlation between the number of intermediate "reasoning" tokens a model generates and its success rate on complex logic tasks. Forcing a model to output its plan before its answer acts as a form of "external memory."

Strategy Implementation Benefit
Chain of Thought "Think step-by-step before answering." Improves logical accuracy.
Sequential Prompting Task A -> Output A -> Task B (using Output A). Reduces "context bloat" and focus loss.
Self-Criticism "Review your previous answer for errors." Reduces hallucinations.

Concrete Example: AIP Agentic Workflow

In AIP Agent Studio, a complex task like "Process an Insurance Claim" is broken down:

  1. Step 1 (Extraction): Extract claimant details from a PDF.
  2. Step 2 (Ontology Lookup): Check the claimant's policy in the Ontology.
  3. Step 3 (Reasoning): Compare the claim against policy limits.
  4. Step 4 (Action): Generate an approval or denial letter.

Constraint Setting and Guardrails

What it is

Constraint Setting involves defining the "negative space" of a prompt—what the model must not do. This includes formatting constraints, length limits, and safety guardrails.

Why it matters

In production systems, "unconstrained" AI is a liability. A model might provide a correct answer but in a format that crashes a downstream UI, or it might inadvertently reveal sensitive information. Constraints ensure the AI behaves as a reliable component of a larger software system.

How it works: Negative Constraints and Formatting

Engineers use explicit "Do Not" instructions and "System Prompts" to enforce boundaries. In Palantir AIP, these are often augmented by Security and Governance features like the Sensitive Data Scanner, which prevents the model from processing PII (Personally Identifiable Information) regardless of the prompt.

Constraint Type Example Instruction Purpose
Negative Constraint "Do not mention competitor names." Brand safety and compliance.
Structural Constraint "Output only valid JSON. No preamble." Programmatic integration.
Length Constraint "Limit the summary to 3 bullet points." UI/UX consistency.
Privacy Constraint "Redact all names and social security numbers." Data security.

Common Pitfalls

  • Prompt Injection: A user attempting to bypass constraints by saying "Ignore all previous instructions."
  • Constraint Conflict: Setting so many constraints that the model has no "room" to generate a valid answer (e.g., "Summarize this book in exactly 5 words without using the letter 'e'").

Integration with Palantir Ontology and AIP Logic

Prompt engineering in AIP is unique because it is "Ontology-aware." Instead of providing the model with raw, disconnected text, AIP allows you to inject Object Sets and Properties directly into the prompt context.

The Role of the Ontology

The Ontology provides the Ground Truth. When you prompt an LLM in AIP Logic, you aren't just asking it to "be smart"; you are asking it to operate on specific, governed data structures. This reduces the need for the model to "know" facts (which leads to hallucinations) and instead focuses its power on "reasoning" over the provided data.

Code Example: AIP Logic Function

This example demonstrates a structured prompt that combines persona, context (from Ontology), and formatting.

# Conceptual AIP Logic Prompt Configuration
def analyze_sensor_data(sensor_object):
    """
    Inputs: sensor_object (Ontology Object: 'Heavy_Machinery_Sensor')
    """
    system_prompt = """
    You are an expert Reliability Engineer. You will receive a JSON representation 
    of a machinery sensor's recent telemetry and its historical maintenance records.
    
    Your task is to:
    1. Identify if the current 'temperature' exceeds the 'threshold' property.
    2. Check the 'last_service_date' to see if it was more than 6 months ago.
    3. Provide a 'risk_score' (0-100).
    
    Constraints:
    - Output ONLY a JSON object.
    - Do not include any conversational filler.
    """
    
    user_prompt = f"""
    Sensor Data: {sensor_object.telemetry_json}
    Maintenance History: {sensor_object.maintenance_history}
    Current Date: {get_current_date()}
    """
    
    # The AIP Logic engine handles the LLM call, 
    # ensuring data governance and zero-data retention.
    return call_llm(system_prompt, user_prompt, model="gpt-4o")

Security, Governance, and Ethics

Responsible AI Principles

Palantir AIP is built on the principle that AI should be Explainable, Traceable, and Reliable. Prompt engineering plays a vital role in this:

  • Explainability: By using Chain-of-Thought prompting, the model's reasoning process is captured in the logs, allowing humans to audit why a decision was made.
  • Traceability: AIP's Distributed Tracing and Logging features record the exact prompt sent to the model, including the versions of the Ontology objects injected into the context.
  • Reliability: Through AIP Evals, teams can prove that a prompt meets safety and accuracy standards before it is deployed to production.

Data Privacy

When using third-party-hosted LLMs (via Foundry's proxy endpoints), AIP ensures Zero Data Retention (ZDR). This means that while your prompt (containing sensitive Ontology data) is sent to the model provider for processing, it is never used to retrain the provider's models and is deleted immediately after the response is generated.

Best Practices for Prompt Engineering - AIP overview - diagram 1
Best Practices for Prompt Engineering - AIP overview - diagram 1

Bring Your Own Model (BYOM)

Key concepts: Registered Models · Function Interfaces · ChatCompletion Interface · TypeScript Functions

How to register and use external or fine-tuned models within the AIP ecosystem.

Bring Your Own Model (BYOM)

In the rapidly evolving landscape of artificial intelligence, the ability to decouple the application layer from the underlying model provider is a critical architectural requirement. While Palantir’s Artificial Intelligence Platform (AIP) provides native access to industry-leading Large Language Models (LLMs) from providers like OpenAI, Anthropic, and Google, enterprise requirements often necessitate the use of custom, fine-tuned, or locally hosted models. The Bring Your Own Model (BYOM) framework allows organizations to integrate these external models into the Palantir ecosystem while maintaining the platform's rigorous standards for security, governance, and observability.

AI_SVGI_SVGThe BYOM Architecture: Illustrating the flow from AIP Logic through the Registered Model abstraction, the TypeScript Function interface, and finally to the external Model Provider via REST API.*

The Philosophy of Model Agnosticism

The core tenet of AIP is model-agnosticism. This means that the higher-level components of the platform—such as AIP Logic, AIP Agent Studio, and Pipeline Builder—do not depend on the specific implementation details of a model. Instead, they interact with a standardized abstraction.

BYOM is not merely a "connector" but a formal integration pattern. It allows a senior engineer to wrap any computational engine (whether it is a fine-tuned Llama-3 instance running on-premises or a specialized legal-domain model) in a way that the platform treats it as a first-class citizen. This ensures that features like AIP Evals and AIP Observability function identically regardless of whether the model is native or external.

Registered Models: The Abstraction Layer

A Registered Model is a metadata entity within the Palantir platform that represents an external model. It acts as a pointer and a configuration set that tells AIP how to communicate with the model and what capabilities it supports.

When you register a model, you are essentially creating a contract. This contract specifies the model's identifier, its input/output constraints, and the logic required to invoke it. This abstraction allows developers to swap models in an AIP Logic workflow without rewriting the underlying prompts or data integrations.

Component Description Role in BYOM
Model Identifier A unique string identifying the model within the enrollment. Routing and referencing in Logic/Agents.
Provider Mapping Links the registered model to a specific Function or REST source. Execution logic.
Capability Set Defines if the model supports streaming, tool calling, or vision. UI enablement in Builder tools.
Georestriction Policy settings defining where data can be sent. Compliance and data sovereignty.

The ChatCompletion Interface

The primary mechanism for integrating a BYOM model is the ChatCompletion Interface. This is a standardized TypeScript interface that mirrors the industry-standard "Chat Completion" pattern (popularized by the OpenAI API but now a de facto standard for generative AI).

To implement BYOM, a developer writes a TypeScript Function that implements this interface. This function acts as a translator: it receives a standardized request from AIP, transforms it into the specific format required by the external model, executes the network call, and then transforms the response back into the AIP-standard format.

Technical Specification of the Interface

The interface typically requires the implementation of a method that handles an array of Message objects. Each message contains a role (system, user, assistant, or tool) and the content.

Definition: The ChatCompletion Contract Let $M$ be the set of all possible messages and $R$ be the set of model responses. A BYOM function $f$ is a mapping $f: M^n \times P \rightarrow R$, where $P$ represents the set of hyperparameters (temperature, top-p, etc.). The function must guarantee that $R$ adheres to the ChatCompletionResponse schema to ensure downstream compatibility with AIP Logic.

Implementation via TypeScript Functions

TypeScript Functions provide the "glue" for BYOM. Because these functions run within the secure, managed environment of Foundry’s sidecars, they can safely access REST API Sources and Webhooks that are configured with sensitive credentials (like API keys or Bearer tokens).

Worked Example: Integrating a Custom LLM

Consider an organization that has deployed a fine-tuned Mistral model on an internal Kubernetes cluster. The model is exposed via a REST endpoint. The following code block demonstrates how a TypeScript function serves as the bridge.

import { ChatCompletionRequest, ChatCompletionResponse, Message } from "@foundry/functions-api";
import { MyExternalModelClient } from "@foundry/external-systems";

export class ModelIntegrationFunctions {
    /**
     * Bridges AIP to a custom-hosted Mistral model.
     * @param request The standardized AIP request object.
     * @returns A promise resolving to the standardized response.
     */
    @Function()
    public async callCustomMistral(request: ChatCompletionRequest): Promise<ChatCompletionResponse> {
        // 1. Transform AIP messages to the external provider's specific format
        const externalPayload = {
            prompt: this.formatPrompt(request.messages),
            max_tokens: request.maxTokens ?? 512,
            temperature: request.temperature ?? 0.7
        };

        try {
            // 2. Execute the call via a pre-configured REST Source
            const response = await MyExternalModelClient.post("/v1/generate", externalPayload);

            // 3. Map the external response back to ChatCompletionResponse
            return {
                choices: [{
                    message: {
                        role: "assistant",
                        content: response.data.text
                    },
                    finishReason: "stop"
                }],
                usage: {
                    promptTokens: response.data.usage.prompt,
                    completionTokens: response.data.usage.completion,
                    totalTokens: response.data.usage.total
                }
            };
        } catch (error) {
            // 4. Handle Rate Limiting specifically for AIP's retry logic
            if (error.status === 429) {
                throw new RateLimitExceededError("External model is throttled.");
            }
            throw error;
        }
    }

    private formatPrompt(messages: Message[]): string {
        // Custom logic to format messages for the specific model's prompt template
        return messages.map(m => `[${m.role}]: ${m.content}`).join("\n");
    }
}

AI_DEMOI_DEMOInteractive Simulation: A visual debugger showing a message entering the TypeScript Function, being transformed into a JSON payload for a REST API, and the subsequent mapping of the response back into the Ontology-compatible format.*

Connectivity: REST API Sources and Webhooks

For the TypeScript function to reach the external model, a REST API Source must be configured in the Palantir Control Panel. This is a critical security step. Instead of hardcoding URLs or credentials in the code, the function references a named source.

  1. Authentication: Supports Basic Auth, API Keys, or OAuth2.
  2. Network Egress: AIP administrators must explicitly whitelist the domain of the external model provider.
  3. Encapsulation: The function environment ensures that the API key is never exposed to the end-user of the AI application.

Operationalization in AIP Logic and Pipeline Builder

Once the TypeScript function is published and the model is registered, it appears in the model selection dropdowns across the platform.

AIP Logic Integration

In AIP Logic, the BYOM model can be used as the primary LLM for a logic block. The platform handles the orchestration, providing the model with access to Ontology objects and tools. Because the BYOM model follows the ChatCompletion interface, AIP Logic can perform "Tool Calling" (Function Calling) by sending the tool definitions to the BYOM function, provided the underlying model supports it.

Pipeline Builder Integration

For batch processing, the BYOM model can be invoked within Pipeline Builder. This allows for large-scale data enrichment (e.g., "Summarize these 1 million support tickets using our fine-tuned internal model"). The platform manages the parallelization and compute allocation, treating the BYOM call as a transformation step in the data pipeline.

Feature Native Model (e.g., GPT-4) BYOM (Custom Model)
Setup Complexity Zero (Toggle in Control Panel) Medium (Requires TS Function)
Customization Low (System Prompts only) High (Fine-tuned weights)
Data Residency Provider-dependent Fully User-controlled
Latency Internet-dependent Network-dependent (can be local)
Cost Compute-seconds / Tokens External Provider Costs + Function Compute

Error Handling and Performance Optimization

A sophisticated BYOM implementation must account for the non-deterministic nature of network communication and model inference.

Rate Limit Propagation

One of the most common pitfalls in BYOM is failing to propagate rate limits. If the external model provider returns a 429 Too Many Requests error, the TypeScript function should catch this and throw a specific RateLimitExceeded error.

Why this matters: AIP has built-in orchestration logic that understands how to handle retries with exponential backoff. If the function simply throws a generic error, AIP might treat it as a permanent failure, causing the entire workflow to crash. By using the correct error type, the developer allows the platform to manage the queueing and retrying of requests efficiently.

Observability and P95 Latency

BYOM models are fully integrated into AIP Observability. This means that every call to the custom model is tracked via Distributed Tracing.

  • Metrics: You can monitor the P95 duration of your custom function calls.
  • Logging: Execution logs from the TypeScript function are available in the Workflow Lineage view.
  • Cost Tracking: While the external model's token costs are managed externally, the compute-seconds used by the TypeScript function to process the request are tracked within AIP.

Security, Governance, and Ethics

Integrating an external model does not exempt the workflow from Palantir’s security and ethical frameworks.

  1. Data Privacy: AIP ensures that data sent to a BYOM model is encrypted in transit. Furthermore, the Sensitive Data Scanner can be used to intercept and redact PII (Personally Identifiable Information) before it leaves the Palantir environment for the external model.
  2. No Retraining Guarantee: Contractual and technical safeguards ensure that data sent to third-party providers via BYOM is not used to retrain their foundational models.
  3. Traceability: Every interaction with the BYOM model is recorded in the audit logs, providing a complete history of what data was sent to which model and what response was received. This is vital for industries with strict Explainability and Transparency requirements.

Common Pitfalls and Best Practices

Pitfall: Inconsistent Token Counting

Different models use different tokenizers (e.g., Tiktoken for OpenAI vs. SentencePiece for Llama). If your BYOM function reports token usage based on a different tokenizer than the model actually uses, your cost tracking and context window management will be inaccurate.

  • Best Practice: Always return the token usage counts provided by the external model's API response rather than calculating them locally in the function.

Pitfall: Timeout Mismatch

AIP functions have a default timeout. If the external model is slow (e.g., generating a very long response), the function might time out before the model finishes.

  • Best Practice: Configure the REST source timeout to match the expected P99 latency of the model, and ensure the TypeScript function is optimized for asynchronous execution.

Pitfall: Over-formatting

Developers often try to "clean" the model's output inside the TypeScript function.

  • Best Practice: Keep the BYOM function as a "thin" pass-through. Let AIP Logic or the application layer handle the parsing and formatting of the response. This keeps the model integration clean and reusable.

Generalization: Beyond Text

While most BYOM use cases focus on ChatCompletion, the framework is extensible. Organizations can bring their own Text Embedding Models using a similar pattern. These models are crucial for RAG (Retrieval-Augmented Generation) workflows, where the custom embedding model ensures that the vector database stays synchronized with the specific semantic nuances of the organization's data.

Theorem of Model Substitution In a well-architected AIP environment, the substitution of a Native Model $M_n$ with a Registered Model $M_{byom}$ should result in zero changes to the downstream Ontology transformations, provided the mapping function $f$ preserves the semantic integrity of the ChatCompletion interface.

AI_FLASHCARDSI_FLASHCARDS Registered Model: A platform-level abstraction representing an external AI model.

  • ChatCompletion Interface: The standard TypeScript contract for exchanging messages and roles with an LLM.
  • RateLimitExceeded: A specific error type that triggers AIP's automated retry logic.
  • REST API Source: A secure configuration for managing external credentials and endpoints.
  • Model-Agnosticism: The design principle that allows AIP tools to work with any underlying model.
  • P95 Duration: A metric used in AIP Observability to track the latency of model responses.

AI_QUIZI_QUIZ. What is the primary purpose of the ChatCompletion interface in a BYOM setup?

  • (A) To train the model on new data.
  • (B) To provide a standardized contract for communication between AIP and the model.
  • (C) To encrypt the data before it leaves Foundry.
  • (D) To bypass the need for a REST API source. Answer: (B)
  1. Why is it critical to throw a RateLimitExceeded error in a custom BYOM function?

    • (A) It reduces the cost of the API call.
    • (B) It allows the platform to perform automated retries with backoff.
    • (C) It prevents the model from being over-trained.
    • (D) It is required by the TypeScript compiler. Answer: (B)
  2. Where are the credentials for an external model stored in a BYOM integration?

    • (A) Hardcoded in the TypeScript Function.
    • (B) In the user's browser local storage.
    • (C) In a REST API Source or Webhook configuration.
    • (D) In the model's system prompt. Answer: (C)
  3. Which tool would you use to monitor the latency of your BYOM model calls?

    • (A) Pipeline Builder.
    • (B) AIP Logic.
    • (C) AIP Observability (Workflow Lineage).
    • (D) Agent Studio. Answer: (C)

AI_STUDY_GUIDEI_STUDY_GUIDE*Key Concepts to Master:**

  • Architectural Flow: Understand the path from a user prompt in AIP Logic, through the Registered Model, into the TypeScript Function, and out to the external REST endpoint.
  • Interface Implementation: Be comfortable with the ChatCompletionRequest and ChatCompletionResponse schemas.
  • Security Model: Explain how Palantir ensures data privacy when communicating with third-party-hosted models.
  • Error Handling: Know the difference between a transient error (like a rate limit) and a permanent error, and how to handle each in code.
  • Integration Points: Identify which AIP features (Logic, Agents, Pipelines) can utilize BYOM and how they benefit from it.
  • Observability: Describe how to use distributed tracing to debug a slow-performing custom model integration.
Bring Your Own Model (BYOM) - AIP overview - diagram 1
Bring Your Own Model (BYOM) - AIP overview - diagram 1

Agentic Workflows with AIP Analyst

Key concepts: Agentic Workflows · Ontology SQL · Analysis Provenance · Workshop Widget · Context Management

Using AIP Analyst for ad-hoc data analysis and autonomous ontology exploration.

Agentic Workflows with AIP Analyst

In the contemporary enterprise landscape, the bottleneck of data-driven decision-making has shifted from data availability to the latency of analysis. Traditional Business Intelligence (BI) models rely on a "request-and-wait" cycle where operational users task data scientists with generating reports. AIP Analyst represents a paradigm shift toward Agentic Workflows, where Large Language Models (LLMs) are not merely passive responders but active participants in the analytical process. By leveraging the Palantir Ontology, AIP Analyst transforms natural language into executable operations, providing a bridge between human intuition and computational scale.

AI_SVGI_SVGThe AIP Analyst Architecture: Illustrating the flow from Natural Language Input through the Agentic Reasoning Loop, interacting with the Ontology via O-SQL, and outputting via Analysis Provenance and Vega Visualizations.*

The Anatomy of Agentic Workflows

An Agentic Workflow is a computational pattern where an AI agent is empowered to autonomously plan, execute, and refine a multi-step sequence of actions to achieve a high-level goal. Unlike a standard chatbot that provides a single-turn response, an agentic system operates in a closed-loop cycle of reasoning and acting.

The ReAct Pattern

AIP Analyst primarily utilizes the ReAct (Reason + Act) framework. When a user submits a query—for example, "Identify the top three flight delays caused by weather in the Northeast and suggest a re-routing strategy"—the agent does not attempt to answer immediately. Instead, it follows a structured internal logic:

  1. Thought: The agent decomposes the request into sub-tasks (e.g., filter flights by region, filter by reason, join with airport metadata).
  2. Action: The agent selects a tool from its repertoire (e.g., search_ontology, execute_sql).
  3. Observation: The agent inspects the output of the action (e.g., a table of delayed flights).
  4. Refinement: If the observation is insufficient, the agent updates its "Thought" and repeats the cycle.

Definition: Agentic Autonomy In the context of AIP, autonomy is defined as the agent's ability to navigate the Ontology—the digital twin of the organization—without explicit hard-coded paths. The agent treats Object Types, Link Types, and Action Types as a dynamic API.

Feature Traditional BI Agentic Analysis (AIP Analyst)
User Input Structured SQL / Drag-and-Drop Natural Language (Ambiguous/Complex)
Logic Definition Pre-defined by Dashboard Author Dynamically generated by LLM
Data Discovery Manual search of tables Autonomous Ontology Discovery
Error Handling Query fails; user debugs Agent observes error; self-corrects
Output Static Charts Interactive, Branching Workflows

Ontology SQL (O-SQL): The Interface of Truth

A critical component of the agent’s ability to interact with data is Ontology SQL (O-SQL). While standard SQL queries raw relational tables, O-SQL is "Ontology-aware." It understands the semantic relationships, security permissions, and object-oriented structure of the Palantir platform.

Why O-SQL Matters

LLMs often struggle with raw database schemas because table names like TBL_TRNS_V2_FINAL lack semantic meaning. AIP Analyst translates the user's intent into O-SQL, which targets the Ontology. Because the Ontology uses human-readable names (e.g., [Flight], [Airport], [Delay Event]), the LLM can generate queries with significantly higher precision and lower hallucination rates.

Technical Implementation

When the agent decides to execute a transformation, it generates an O-SQL block. This block is not just a string but a tracked resource that respects Analysis Provenance.

-- Example of an Agent-generated O-SQL query for flight analysis
SELECT 
    f.flight_number, 
    f.departure_delay, 
    a.city, 
    f.weather_condition
FROM 
    `aviation.Flight` f
JOIN 
    `aviation.Airport` a ON f.origin_airport_id = a.id
WHERE 
    a.region = 'Northeast' 
    AND f.weather_condition IS NOT NULL
ORDER BY 
    f.departure_delay DESC
LIMIT 3;

Variations in Object Transformation

The agent can perform several types of transformations:

  • Object Set Filtering: Narrowing down a collection of objects based on properties.
  • Link Traversal: Moving from a [Customer] object to their [Orders] via defined relationships.
  • Aggregations: Calculating P95 durations or sum of costs across an object set.

Analysis Provenance: The "How" and "Why"

In enterprise environments, an answer is only as good as its audit trail. Analysis Provenance is the mechanism by which AIP Analyst records every step of its reasoning and data transformation.

The Provenance Graph

Every result generated by the Analyst is backed by a Directed Acyclic Graph (DAG). Each node in this graph represents a specific state of the analysis:

  • Input Nodes: The raw natural language prompt.
  • Reasoning Nodes: The LLM’s internal "Thought" process.
  • Action Nodes: The specific O-SQL queries or tool calls executed.
  • Data Nodes: The resulting object sets or dataframes.

This transparency solves the "Black Box" problem of AI. A user can click on any chart or number and see exactly which filters were applied and which objects were included.

Provenance Component Description Metadata Captured
Logic Trace The step-by-step reasoning chain. Model ID, Temperature, Prompt Version
Data Lineage The path from raw Object to Result. Object RID, Branch ID, Transaction Time
Execution History The physical compute details. Compute-seconds, Token count, Latency
Verification Human-in-the-loop validation. User ID, Timestamp of Approval

AI_DEMOI_DEMOInteractive Simulation: A user enters "Show me high-risk suppliers." The demo visualizes the Agentic Loop: 1. Search Ontology for 'Supplier' -> 2. Identify 'Risk Score' property -> 3. Generate O-SQL for Score > 80 -> 4. Render results as a Vega-Lite bar chart.*

Context Management and Branching

One of the most sophisticated features of AIP Analyst is its ability to manage Context and Branching. In a complex investigation, an analyst rarely follows a linear path. They might explore a hypothesis, find it's a dead end, and want to return to a previous state.

Statefulness in Conversations

AIP Analyst maintains a session state that includes the current "Object Set" in focus. This allows for follow-up questions like "Now show only the ones in Texas." The agent understands that "the ones" refers to the result of the previous turn.

Parallel Analysis Paths (Branching)

AIP Analyst allows users to create Branches in the analysis. This is analogous to git branch for data investigation.

  1. Main Path: Analyzing general sales trends.
  2. Branch A: Investigating the impact of a specific marketing campaign.
  3. Branch B: Investigating the impact of a competitor's price drop.

Users can toggle between these branches, comparing the resulting visualizations and provenance graphs side-by-side. This is essential for Scenario Analysis and Root Cause Analysis.

The Workshop Widget: Embedding Agency

While AIP Analyst exists as a standalone interface, its true power is realized when embedded into operational applications via the AIP Analyst Workshop Widget.

Integration Mechanics

Developers can drag the Analyst widget into a Palantir Workshop module. This allows for a "Hybrid" UI where structured components (tables, maps) live alongside an agentic chat interface.

Pre-loading Context

The widget can be configured with Initial Context. For example, if a user is looking at a specific "Production Plant" in a dashboard, the Analyst widget can be pre-loaded with that specific object. When the user asks "What are the current bottlenecks?", the agent already knows the scope is limited to that plant.

Parameter Purpose Example Value
Initial Object Set Limits the agent's initial search space. [Current_Selected_Plant]
Available Tools Restricts which actions the agent can take. [Search, SQL, Write_Back]
System Prompt Defines the persona and constraints. "You are a supply chain expert..."
Variable Binding Connects UI variables to the Agent's memory. Selected_Date_Range

Observability and Performance Monitoring

Deploying an agentic workflow requires rigorous monitoring to ensure reliability and cost-efficiency. AIP provides a suite of Observability tools specifically for these workflows.

Metrics and P95 Duration

Because LLMs are non-deterministic, execution times can vary. AIP tracks the P95 duration (the time within which 95% of requests complete) for agentic steps. If the agent enters an infinite loop or a particularly complex "Reasoning" phase, observability alerts can trigger.

Compute Usage (Tokens and Seconds)

AIP compute usage is measured in compute-seconds and tokens.

  • Input Tokens: The context provided to the agent (Ontology metadata, history).
  • Output Tokens: The reasoning and O-SQL generated.
  • Compute-Seconds: The actual time the LLM provider spends processing the request.

Pitfall: Context Window Saturation A common mistake is providing too much metadata in the initial context. If the agent is aware of 10,000 object types, the prompt becomes bloated, increasing cost and decreasing accuracy. Ontology Resource Discovery should be used to dynamically prune the metadata sent to the LLM based on the user's query.

Security, Governance, and Ethics

AIP Analyst operates within the strict security boundaries of the Palantir platform. It is not a "God-mode" tool; it is a "User-mode" tool powered by AI.

  1. Mandatory Access Control (MAC): If a user does not have permission to see "Salary" data in the Ontology, the AIP Analyst cannot see it either, even if the LLM "wants" to use it for an analysis.
  2. Model Retraining Policy: Data sent to third-party LLM providers (like OpenAI or Anthropic) via AIP is never used to retrain their base models. It is processed in transient memory and deleted.
  3. Explainability: Through Analysis Provenance, every decision made by the agent is auditable. This fulfills the Transparency requirement of Responsible AI.

Common Pitfalls in Agentic Workflows

  • Ambiguous Ontology Naming: If two objects are named Client and Customer, the agent may struggle to choose the correct one. Clean, semantic naming in the Ontology is the "Prompt Engineering" of the data layer.
  • Over-reliance on SQL: Sometimes a simple link traversal is more efficient than a complex O-SQL join. High-performing agents are tuned to prefer the simplest tool for the job.
  • Ignoring Evals: Failing to use AIP Evals to test the agent against a "Golden Set" of questions often leads to regressions when the underlying LLM is updated.

AI_FLASHCARDSI_FLASHCARDS Agentic Workflow: A multi-step loop of reasoning and action (ReAct) to solve complex tasks.

  • O-SQL: Ontology-aware SQL that understands semantic object relationships.
  • Analysis Provenance: The auditable DAG showing the logic and data behind an AI result.
  • Workshop Widget: An embeddable component to bring agentic analysis into custom apps.
  • Context Branching: The ability to create parallel, non-linear analysis paths.
  • Compute-Seconds: The metric used to track LLM processing time and cost.

AI_QUIZI_QUIZ. How does AIP Analyst handle a query it cannot answer in a single step?

  • A) It returns an error.
  • B) It uses the ReAct pattern to reason, act, and observe in a loop.
  • C) It asks the user to write the SQL manually.
  • Answer: B
  1. What is the primary benefit of O-SQL over standard SQL in an agentic context?

    • A) It is faster to execute.
    • B) It uses semantic names from the Ontology, making it easier for LLMs to generate accurately.
    • C) It bypasses security permissions.
    • Answer: B
  2. In Analysis Provenance, what does a "Data Node" represent?

    • A) The LLM's internal thought.
    • B) The specific version of the model used.
    • C) The resulting object set or dataframe from an action.
    • Answer: C
  3. True or False: AIP Analyst can see data that the logged-in user is not authorized to view.

    • Answer: False. It respects all platform-level security and access controls.

AI_STUDY_GUIDEI_STUDY_GUIDE*Key Learning Objectives:**

  1. Understand the ReAct Framework: Be able to explain how an agent moves from "Thought" to "Action" to "Observation."
  2. Master Ontology Integration: Explain why the Ontology is the prerequisite for effective agentic workflows.
  3. Provenance as Trust: Articulate how DAGs and execution history provide the transparency required for enterprise AI.
  4. Application Design: Learn how to use the Workshop Widget to create "AI-augmented" operational tools.
  5. Efficiency and Governance: Identify the metrics (tokens, compute-seconds) and security protocols (MAC, Retraining policies) that govern AIP Analyst.

Further Reading:

  • AIP Logic vs. AIP Analyst: Compare the developer-centric logic builder with the analyst-centric ad-hoc tool.
  • Vega-Lite Documentation: Understand the grammar of graphics used by the Analyst to render visualizations.
  • Palantir Learning Portal: Complete the "AIP Workflow Speedrun" for hands-on experience.
Agentic Workflows with AIP Analyst - AIP overview - diagram 1
Agentic Workflows with AIP Analyst - AIP overview - diagram 1

Testing and Benchmarking with AIP Evals

Key concepts: Evaluation Suite · Target Function · Grid Search · Ontology Simulation · LLM Trace Viewer

A comprehensive framework for evaluating LLM performance and ensuring the reliability of AI functions.

Testing and Benchmarking with AIP Evals

In the paradigm of classical software engineering, testing is largely deterministic. For a given input $x$, a function $f$ is expected to yield a consistent output $y$. However, the integration of Large Language Models (LLMs) into operational workflows introduces a fundamental challenge: stochasticity. Because LLMs operate on probabilistic token prediction, the same prompt can yield subtly or significantly different results across multiple invocations.

AIP Evals is Palantir’s rigorous framework designed to transition LLM development from "vibe-based" prompting to empirical engineering. It provides the infrastructure to benchmark Target Functions—whether they are AIP Logic flows, Agentic functions, or raw code—against structured Evaluation Suites. By quantifying performance through systematic iteration, AIP Evals ensures that AI-driven automation remains reliable, safe, and aligned with enterprise requirements.

AI_SVGI_SVG## The Evaluation Suite and the Target Function

At the heart of the benchmarking process lies the relationship between the Target Function and the Evaluation Suite. To understand this, we must first define the mathematical objective of an evaluation: to minimize the delta between the model's output and a defined "ground truth" or "ideal behavior."

1. What it is

A Target Function is the specific unit of logic being tested. In AIP, this is typically an AIP Logic function (a no-code LLM workflow), an AIP Agent function (a tool-use capability), or a code-authored function.

An Evaluation Suite is a curated collection of Test Cases. Each test case consists of:

  • Inputs: The variables passed to the Target Function.
  • Expected Output (Optional): The "ground truth" used for comparison.
  • Evaluation Functions: Logic (often another LLM or a deterministic script) that scores the Target Function's output.

2. Why it matters

Without a structured suite, developers often fall into the trap of "anecdotal testing"—changing a prompt and checking if it works for one specific case, only to unknowingly break five others. The Evaluation Suite provides a regression testing framework, ensuring that optimizations in one area do not cause regressions elsewhere.

3. How it works: The Evaluation Pipeline

The process follows a discrete sequence:

  1. Selection: The user selects a Target Function version.
  2. Execution: The Target Function is run against every input in the Evaluation Suite.
  3. Scoring: The outputs are passed to Evaluation Functions (e.g., "Does this output contain PII?", "Is the sentiment positive?", "Does the JSON match the schema?").
  4. Aggregation: Results are compiled into a dashboard showing pass/fail rates, mean scores, and latency.
Component Responsibility Example
Target Function The "Student" being tested generate_supply_chain_summary(region)
Test Case The "Exam Question" region = "North America"
Evaluation Function The "Grader" check_contains_keyword(output, "Shortage")
Metric The "Final Grade" % of cases where "Shortage" was correctly identified

4. Common Pitfalls

  • Overfitting to Test Cases: Designing a prompt that works perfectly for the 10 cases in your suite but fails on the 11th. Solution: Ensure test cases represent the full variance of real-world data.
  • Weak Evaluation Functions: Using an LLM grader that is too "lenient." Solution: Use deterministic checks (Regex, Schema validation) alongside LLM-based grading.

Grid Search: Hyperparameter Optimization for Prompts

In traditional machine learning, a Grid Search is used to find the optimal combination of hyperparameters (like learning rate or dropout). In AIP Evals, Grid Search is applied to the "LLM Configuration Space."

1. What it is

Grid Search in AIP Evals is the systematic execution of a Target Function across a matrix of different Models and Prompts. It allows a developer to answer the question: "Which specific combination of Model X and Prompt Version Y yields the highest accuracy at the lowest cost?"

2. Why it matters

The "best" model is not always the largest one. A smaller, cheaper model (like GPT-3.5 or Claude Haiku) might perform just as well as a frontier model (GPT-4 or Claude Opus) for a specific, narrow task if the prompt is well-engineered. Grid Search provides the empirical data needed to make these cost-performance trade-offs.

3. Mechanics of the Grid

Imagine a scenario where you have 3 prompt variations and 3 candidate models. A Grid Search will run $3 \times 3 = 9$ distinct "configurations" against your entire test suite.

Theorem of Prompt Sensitivity: Small perturbations in prompt structure (e.g., changing "Summarize this" to "Provide a concise executive summary of this") can lead to non-linear changes in output quality across different model architectures.

4. Concrete Example: Optimizing a Classifier

Suppose we are building a function to classify customer support tickets into "Urgent" or "Routine."

# Conceptual representation of a Grid Search Configuration
search_space = {
    "models": ["gpt-4o", "claude-3-sonnet", "llama-3-70b"],
    "prompts": [
        "Classify this ticket: {{input}}",
        "You are a support agent. Classify the urgency of: {{input}}. Reply only with the label.",
        "Few-shot example: [Ticket: My house is on fire] -> Urgent. Now classify: {{input}}"
    ]
}
Configuration Accuracy Avg. Latency Cost per 1k runs
GPT-4o + Few-shot 98% 1.2s $15.00
Claude-3 + Role-play 94% 0.8s $3.00
Llama-3 + Basic 82% 0.5s $0.50

Analysis: If the business requirement is >90% accuracy, the Claude-3 configuration is the "winner" as it is 5x cheaper than GPT-4o while meeting the threshold.

AI_DEMO--

Ontology Simulation: Safe Side-Effect Testing

One of the most powerful features of Palantir AIP is its ability to not just "chat," but to act—specifically by editing the Ontology (e.g., updating an order status, creating a new sensor log). Testing these actions in a production environment is dangerous.

1. What it is

Ontology Simulation is a sandboxing mechanism within AIP Evals. It allows an LLM or Agent to perform "writes" to the Ontology during a test run without actually persisting those changes to the real database.

2. Why it matters

If an Agent is designed to "Cancel all overdue orders," a bug in the prompt might lead it to "Cancel all orders." In a standard testing environment, this would be catastrophic. Ontology Simulation provides a "Copy-on-Write" layer where the Agent interacts with a virtual snapshot of the data.

3. How it works

  1. Snapshotting: When a test starts, AIP Evals creates a temporary session context.
  2. Interception: Any write or update command issued by the Target Function is intercepted by the simulation layer.
  3. Virtual State: The simulation maintains a "delta" of changes. If the Agent creates "Object A," and then performs a search, "Object A" will appear in the search results within that specific test session, even though it doesn't exist in the real Ontology.
  4. Validation: Evaluation functions can then query this virtual state to verify if the correct edits were made.
Feature Production Ontology Ontology Simulation
Data Persistence Permanent Ephemeral (Deleted after test)
Visibility All Users Only the Eval Session
Risk of Corruption High Zero
State Awareness Real-time Supports "What-if" scenarios

4. Common Pitfalls

  • Simulation Lag: If the production Ontology is massive, the simulation may only include a subset of objects. If the Agent relies on an object not in the subset, the test will fail. Solution: Use "Object Set" definitions to explicitly include required context in the test case.

The LLM Trace Viewer: Debugging the Reasoning Chain

When a complex Agent fails a test, simply knowing "it gave the wrong answer" is insufficient. We need to know why. Was it a retrieval failure? A logic error? A tool-use hallucination?

1. What it is

The LLM Trace Viewer is a distributed tracing tool optimized for agentic workflows. It provides a chronological, hierarchical view of every step the AI took to reach a conclusion.

2. Why it matters

Modern AI applications are rarely a single prompt. They are "chains" or "graphs" of multiple calls. The Trace Viewer allows developers to perform "Root Cause Analysis" (RCA) on non-deterministic failures.

3. How it works: The Anatomy of a Trace

A trace is composed of Spans. Each span represents a discrete unit of work.

  • Input Span: The initial user query.
  • Retrieval Span: The query sent to the Vector Store/Ontology and the results returned.
  • Reasoning Span: The LLM's "Thought" process (often hidden from the end-user but visible in traces).
  • Tool Call Span: The specific parameters sent to a function (e.g., get_weather(city="London")).
  • Output Span: The final response.

4. Concrete Example: Debugging a "Hallucination"

  • Test Result: FAIL (Agent said "Part 123 is in stock," but it is actually "Out of Stock").
  • Trace Analysis:
    1. Span 1 (Input): "Is Part 123 available?"
    2. Span 2 (Tool Call): Agent calls get_inventory(part_id="123").
    3. Span 3 (Tool Output): Function returns status: "Out of Stock".
    4. Span 4 (LLM Reasoning): "The tool says out of stock, but I should check if there's a substitute... actually, I'll just say it's in stock." (Error found here!)
  • Fix: Adjust the system prompt to enforce strict adherence to tool outputs.

Metrics and Benchmarking: Quantitative Rigor

Final evaluation results are distilled into metrics. In AIP Evals, these are categorized into Deterministic and Heuristic metrics.

1. Deterministic Metrics

These are binary or mathematical checks that do not require an LLM to judge.

  • JSON Validity: Did the model return valid JSON?
  • Latency (P95): Does the function return a result within 2 seconds in 95% of cases?
  • Token Usage: How many input/output tokens were consumed (cost tracking)?

2. Heuristic (LLM-as-a-Judge) Metrics

These use a "Grader LLM" to evaluate qualitative aspects.

  • Faithfulness: Does the answer only use information provided in the context?
  • Relevance: Does the answer actually address the user's question?
  • Tone/Style: Does the response follow the brand guidelines (e.g., "Professional," "Concise")?
Metric Type Example Best For
Exact Match output == "Yes" Classification, Boolean logic
Regex matches(r"\d{3}-\d{2}-\d{4}") Format validation (SSN, Phone)
Semantic Similarity cosine_similarity(output, target) > 0.9 Paraphrasing, Summarization
LLM Rubric "Rate helpfulness from 1-5" Open-ended chat, Creative writing

3. The "Evaluation Loop"

The goal of AIP Evals is to create a flywheel:

  1. Run Evals on the current version.
  2. Identify Weaknesses via the Trace Viewer.
  3. Hypothesize Fixes (e.g., "I need to add a few-shot example for edge case X").
  4. Implement Fix in AIP Logic/Agent Studio.
  5. Re-run Evals to confirm improvement and ensure no regressions.

Testing and Benchmarking with AIP Evals - AIP overview - diagram 1
Testing and Benchmarking with AIP Evals - AIP overview - diagram 1

AIP Observability and Monitoring

Key concepts: Workflow Lineage · Distributed Tracing · P95 Duration · Execution History

Monitoring execution history, performance metrics, and resource consumption across AIP.

AIP Observability and Monitoring

In the context of Palantir’s Artificial Intelligence Platform (AIP), observability is not merely a post-deployment convenience; it is a fundamental requirement for operationalizing non-deterministic models within a deterministic enterprise environment. Unlike traditional software systems where a specific input predictably yields a specific output, AI-integrated workflows introduce stochastic variance. AIP Observability provides the telemetry, visualization, and analytical tools necessary to transform these "black box" AI interactions into transparent, auditable, and optimizable business processes.

Integrated directly within Workflow Lineage, AIP Observability allows engineers and operators to track the lifecycle of an AI decision—from the initial prompt and retrieval of Ontology objects to the final model output and subsequent action. By synthesizing metrics, distributed tracing, and execution history, the platform ensures that AI remains accountable, performant, and cost-effective.

AI_SVGI_SVG## The Architecture of Visibility: Workflow Lineage

Workflow Lineage serves as the foundational map for all AIP operations. In a complex enterprise ecosystem, an AI agent rarely acts in isolation. It queries the Ontology, interacts with legacy databases via Apollo-managed pipelines, and triggers actions in external ERP systems. Workflow Lineage captures these interdependencies, providing a directed acyclic graph (DAG) of how data and logic flow through the system.

What it is

Workflow Lineage is a structural representation of the end-to-end data and logic path. It connects the "upstream" data sources (Foundry datasets, Ontology objects) to the "downstream" AI consumers (AIP Logic functions, Agent Studio assistants).

Why it matters

Without lineage, debugging an AI failure is nearly impossible. If an agent provides an incorrect answer, the root cause could be a hallucination, a poorly phrased prompt, or—most critically—stale data in the underlying Ontology. Lineage allows developers to "travel back in time" to see exactly what state the data was in when the model made a specific decision.

Mechanics of Lineage Tracking

  1. Object-Level Provenance: Every object retrieved by an LLM is tagged with its source metadata.
  2. Logic Versioning: AIP tracks which version of an AIP Logic function or prompt template was executed.
  3. Action Attribution: When an agent performs a "write-back" to the Ontology, that change is linked to the specific execution ID of the agent.
Feature Description Importance in AI Workflows
Provenance Tracking the origin of data used in a prompt. Prevents "garbage in, garbage out" by identifying stale data.
Versioning Recording the exact model and prompt version used. Essential for A/B testing and regression analysis.
State Capture Saving the context window at the moment of execution. Critical for auditing and reproducing non-deterministic errors.

Distributed Tracing in Agentic Workflows

As AI systems move from simple chat interfaces to Agentic Workflows—where an LLM might call multiple tools, perform sub-queries, and iterate on its own reasoning—traditional logging becomes insufficient. Distributed Tracing provides a granular view of these multi-step processes.

What it is

Distributed Tracing is a method used to monitor applications, especially those built on microservices or multi-step agentic architectures. In AIP, a Trace represents the entire journey of a single user request, while a Span represents a single unit of work within that trace (e.g., an LLM call, an Ontology search, or a Python function execution).

How it Works: The Trace Hierarchy

When an AIP Agent receives a query, a unique Trace ID is generated. As the agent decomposes the query into tasks:

  • Root Span: The overall agent execution.
  • Child Spans: Individual steps, such as "Search for Aircraft Parts" or "Summarize Maintenance Logs."
  • Metadata: Each span captures input tokens, output tokens, latency, and model parameters (e.g., temperature).

Key Insight: Distributed tracing allows for the identification of "bottleneck steps." If an agent takes 30 seconds to respond, tracing reveals whether the delay was caused by a slow LLM response, a complex SQL query in the Ontology, or a high-latency external API call.

AI_DEMOI_DEMO### Implementation Example: Trace Structure The following conceptual structure illustrates how a trace captures a multi-step AIP Logic execution.

{
  "trace_id": "aip-8823-xf21",
  "root_span": {
    "operation": "Agent_Customer_Support",
    "duration_ms": 4500,
    "spans": [
      {
        "id": "span-1",
        "operation": "Ontology_Search",
        "query": "SELECT * FROM [Customers] WHERE id = '123'",
        "duration_ms": 200
      },
      {
        "id": "span-2",
        "operation": "LLM_Inference",
        "model": "gpt-4-o",
        "input_tokens": 1200,
        "output_tokens": 150,
        "duration_ms": 3800,
        "cost_compute_seconds": 0.45
      },
      {
        "id": "span-3",
        "operation": "Action_Writeback",
        "target": "Support_Tickets",
        "duration_ms": 500
      }
    ]
  }
}

Performance Metrics and P95 Duration

In observability, averages are often misleading. If 90% of AI responses take 2 seconds, but 10% take 60 seconds, the "average" latency of 7.8 seconds doesn't accurately reflect the experience of either group. This is why AIP focuses on P95 Duration.

Defining P95 and Tail Latency

  • P50 (Median): The duration that 50% of executions are below.
  • P95: The duration that 95% of executions are below. This represents the "worst-case" scenario for the vast majority of users.
  • P99: The duration that 99% of executions are below, often capturing extreme outliers like network timeouts or cold starts.

Why P95 Matters for LLMs

LLM latency is highly variable. It depends on:

  1. Prompt Length: More input tokens increase processing time.
  2. Generation Length: LLMs generate text token-by-token; longer responses take linearly more time.
  3. Queueing: High demand on model providers can lead to "time-to-first-token" (TTFT) delays.

By monitoring P95, engineers can set realistic Service Level Objectives (SLOs). If the P95 duration exceeds a threshold (e.g., 15 seconds), it may trigger an optimization effort, such as switching to a smaller, faster model (e.g., GPT-4o-mini) or implementing prompt caching.

Metric Definition Target for AI Agents
TTFT Time to First Token (the start of the response). < 1.5 seconds
TPS Tokens Per Second (generation speed). > 30 tokens/sec
P95 Duration Total time for 95% of requests. < 10 seconds
Error Rate Percentage of failed model calls or tool timeouts. < 1%

Execution History and 30-Day Retention

AIP maintains a detailed Execution History, typically spanning a rolling 30-day window. This history is the primary interface for debugging and resource optimization.

What it is

Execution History is a chronological log of every invocation of an AIP resource. It includes the status (Success/Failure), the user who triggered it, the duration, and the resource consumption.

The Role of Resource Usage and Cost Tracking

AIP compute usage is fundamentally different from traditional CPU/RAM monitoring. It is measured in compute-seconds, which are derived from the number of tokens processed and the specific model used.

  • Input Tokens: The context provided to the model (instructions + data).
  • Output Tokens: The response generated by the model.
  • Regional Rates: Costs vary based on the cloud region and the provider (OpenAI, Anthropic, Google, etc.).

Common Pitfalls in Execution History Analysis

  1. Ignoring Token Bloat: Developers often include massive amounts of unnecessary data in the prompt context. Execution history reveals when input tokens are disproportionately high compared to the utility of the output.
  2. Overlooking Retries: If an agent is configured to retry on failure, the execution history might show a "Success," but the distributed trace will reveal multiple costly failed attempts hidden within that single success.
  3. Context Window Overflows: History logs will flag when a workflow exceeds the model's maximum context window, leading to truncated data and degraded performance.

Integration with AIP Evals

Observability is the "read" side of the AI lifecycle; AIP Evals is the "test" side. The two are inextricably linked. When a developer identifies a problematic execution in the history (e.g., a P95 outlier or a logic error), they can promote that specific execution to a Test Case in AIP Evals.

The Feedback Loop

  1. Observe: Identify a failure in the 30-day Execution History.
  2. Isolate: Use Distributed Tracing to find the exact span where the logic failed.
  3. Evaluate: Import the inputs/outputs into AIP Evals.
  4. Iterate: Modify the prompt or model, and run the Eval suite to ensure the fix works and doesn't cause regressions in other areas.

Security, Governance, and Transparency

Observability in AIP is governed by strict security protocols to ensure that monitoring doesn't compromise data privacy.

Zero Data Retention (ZDR) and Logging

Palantir ensures that while logs are available for debugging within the Foundry/AIP environment, customer data is not used by third-party model providers for retraining. The Sensitive Data Scanner can be integrated into the observability pipeline to redact PII (Personally Identifiable Information) before it is stored in execution logs or sent to an LLM.

Auditability

Because AIP records the Workflow Lineage, organizations can provide a full audit trail for AI-driven decisions. If a regulatory body asks why a specific loan was denied or a specific maintenance task was prioritized, the organization can produce the exact trace, the data used from the Ontology, and the model's reasoning steps.

Theorem of AI Transparency: The trustworthiness of an AI system is directly proportional to the granularity of its observability and the permanence of its lineage.

Summary of Monitoring Perspectives

Different stakeholders within an organization utilize AIP Observability through different lenses:

Stakeholder Primary Tool Key Metric
Data Engineer Workflow Lineage Data Freshness, Pipeline Health
AI Developer Distributed Tracing, AIP Evals Latency, Accuracy, Token Efficiency
Operations Manager Execution History Success Rate, Task Completion Time
FinOps / Admin Resource Usage Tracking Compute-seconds, Budget vs. Actual
AIP Observability and Monitoring - AIP overview - diagram 1
AIP Observability and Monitoring - AIP overview - diagram 1

Source Materials

Study AIP overview with AI — Free on Lykke

Sign up for free to generate personalized flashcards, quizzes, and study guides from this course. Chat with an AI tutor that knows the material.

Get Started Free

View this course wiki on Lykke · Browse all public course wikis