Ontology building

Institution: MIT

View original course

24 study materials · 6 sections

This course provides a comprehensive guide to building and managing the Palantir Ontology, the operational layer that transforms raw data into a digital twin of an organization. Students will explore the foundational semantic and kinetic elements that enable complex business logic and secure data orchestration. The curriculum covers everything from initial design best practices and administrative management to advanced topics like AI/ML operationalization and semantic search workflows.

Course Sections

Foundations of the Palantir Ontology

Key concepts: Digital twin · Semantic elements (Objects, Properties, Links) · Kinetic elements (Actions, Functions) · Operational layer · Connectivity at scale

Introduction to the core concepts of the Ontology as an operational layer and digital twin.

Foundations of the Palantir Ontology

The Palantir Ontology represents a paradigm shift in enterprise data architecture. While traditional data management focuses on the storage and retrieval of records within siloed databases, the Ontology serves as a dynamic operational layer—a digital twin of the entire organization. It is the connective tissue that sits between raw data assets (datasets, streams, and models) and the end-users who make decisions.

In this lecture, we will explore the foundational components of the Ontology, moving from the static semantic definitions of real-world entities to the kinetic elements that drive business logic and operational change.

[AI_INFOGRAPHIC: The Architecture of the Palantir Ontology. This visual should depict three layers: 1) The Data Foundation (Raw Datasets, Models, Media), 2) The Ontology Layer (Objects, Links, Actions, Functions), and 3) The Application Layer (Workshop, Quiver, Carbon). Arrows should show the bi-directional flow of data and writeback.]

The Operational Layer and the Digital Twin

At its core, the Ontology is a digital twin. In systems engineering, a digital twin is a virtual representation that serves as the real-time digital counterpart of a physical object or process. In the context of Palantir, this definition expands to include abstract concepts like "Legal Cases," "Financial Transactions," or "Supply Chain Routes."

What it is

The Ontology is a semantic mapping of integrated data assets to real-world entities. Mathematically, we can view the Ontology as a directed multigraph $G = (V, E)$, where:

  • $V$ is the set of Object Types (vertices).
  • $E$ is the set of Link Types (edges) connecting these objects.
  • Each $v \in V$ possesses a set of attributes $A$, known as Properties.

Why it matters

Traditional data architectures suffer from "semantic drift," where the meaning of data is lost as it moves through various ETL (Extract, Transform, Load) pipelines. By the time a business user sees a column named TXN_ID_04, they lack the context to know if this represents a completed sale or a pending authorization. The Ontology solves this by providing a shared source of truth. It translates technical metadata into a common language, enabling connectivity at scale.

How it works: The Mapping Function

The transition from raw data to the Ontology involves a mapping function $M: D \rightarrow O$, where $D$ is the domain of raw data rows and $O$ is the range of Ontological objects. This is not a simple renaming; it involves:

  1. Primary Key Mapping: Ensuring each object has a unique, stable identifier (the "Object RID").
  2. Type Casting: Converting raw strings or integers into semantically meaningful types (e.g., Geopoints, Timestamps).
  3. Standardization: Applying logic to ensure that a "Status" property across different source systems is normalized to a common set of values.

Semantic Elements: The "Nouns" of the Organization

Semantic elements define the structure and meaning of the data. They are the static components that describe what the organization is.

Object Types

An Object Type is a schema for a real-world entity. If you think of a dataset as a spreadsheet, an Object Type is the definition of what a row in that spreadsheet represents.

Properties

Properties are the individual fields or attributes of an Object Type. Unlike standard database columns, properties in the Ontology carry additional metadata, including:

  • Visibility: Who can see this specific attribute?
  • Formatting: How should a currency or date be displayed to the user?
  • Base Type vs. Semantic Type: A property might be a String (base type) but represent a Flight Number (semantic type).
Feature Database Column Ontology Property
Data Type Primitive (Int, String, etc.) Semantic (Geopoint, User, Date)
Metadata Minimal (Nullable, Length) Rich (Display hints, Units, PII markers)
Access Control Table-level Property-level (via Mandatory Labeling)
Searchability Keyword/Index Semantic/Vector-based

Link Types

Link Types define the relationships between objects. They are the "verbs" that connect nouns in a static sense (e.g., a "Pilot" flies an "Aircraft"). Links are critical for graph-based analysis, allowing users to traverse the organization’s data.

Link Cardinality is a vital concept here:

  • 1:1 (One-to-One): An Employee has one Work Laptop.
  • 1:M (One-to-Many): A Department has many Employees.
  • M:M (Many-to-Many): A Student attends many Classes, and a Class has many Students.

Key Insight: Links are not just foreign keys. In the Palantir Ontology, links are first-class citizens that enable "Object Traversal." This allows an analyst to start at a "Shipment" object and instantly jump to the "Warehouse" it originated from, and then to the "Manager" of that warehouse, without writing a single SQL JOIN.


Kinetic Elements: The "Verbs" of the Organization

A digital twin that cannot move is merely a map. To be an operational layer, the Ontology must include Kinetic Elements—logic that allows users to interact with and change the state of the world.

Action Types

Action Types encapsulate business logic. They define how data can be modified, ensuring that any "writeback" to the system follows strict rules.

  1. Validation: Before a user can "Update Flight Status," the Action Type checks if the user has the required permissions and if the new status is a valid transition (e.g., you cannot move a flight from "Arrived" back to "In Air").
  2. Side Effects: When an action is taken, it can trigger multiple changes. Changing a "Work Order" to "Complete" might automatically update the "Inventory" count of parts used.
  3. Submission Criteria: These are the "guardrails." For example, a "Credit Limit Increase" action might require a justification text of at least 50 characters and an attachment.

Functions

Functions are blocks of TypeScript code that run on the Ontology. They provide the computational power to perform complex logic that exceeds simple property mapping.

// Example: A Function to calculate the "Risk Score" of a Supplier
// This function traverses links to find all delayed shipments.

import { Function, Integer, Objects } from "@foundry/functions-api";
import { Supplier } from "@foundry/ontology-api";

export class SupplierLogic {
    @Function()
    public async calculateRiskScore(supplier: Supplier): Promise<Integer> {
        const delayedShipments = await supplier.shipments
            .filter(s => s.status.exactMatch("DELAYED"))
            .all();
        
        // Logic: Base risk is 0. Add 10 points for every delayed shipment.
        // Cap the risk score at 100.
        let score = delayedShipments.length * 10;
        return Math.min(score, 100);
    }
}

Why Kinetic Elements Matter

Without Actions and Functions, the Ontology is read-only. Kinetic elements enable Decision Capture. When a user clicks a button in an application to "Approve Loan," that decision is captured as an Action, the Ontology is updated, and the underlying data systems are notified via Data Writeback. This creates a feedback loop where the digital twin stays in sync with the physical world.


Connectivity at Scale: Interfaces and Polymorphism

As an organization grows, its Ontology can become massive, potentially containing thousands of Object Types. To maintain usability and prevent "The God Object" (an anti-pattern where one object tries to do everything), Palantir utilizes Interfaces and Object Polymorphism.

Interfaces

An Interface is a shared definition that multiple Object Types can implement. For example, an organization might have Truck, Airplane, and Cargo Ship object types. While they are different, they all share common traits of a "Vehicle."

By creating a Vehicle Interface, an application developer can build a "Fuel Tracking Dashboard" that works for any object implementing the Vehicle interface, regardless of its specific type.

Object Polymorphism

This allows for "plug-and-play" logic. If a Function is written to calculate the "Maintenance Cost" of a Vehicle interface, it will automatically work for any new object type (like Drone) added to the Ontology later, provided that Drone implements the Vehicle interface.

Concept Definition Purpose
Object Type Concrete implementation (e.g., Boeing 747) Specificity and data backing.
Interface Abstract definition (e.g., Aircraft) Interoperability and abstraction.
Shared Property A property used across types (e.g., Serial Number) Consistency in naming and data types.

Semantic Search and Operational AI

The modern Ontology is not just a structured graph; it is an AI-ready asset. By integrating Semantic Search and Model Integration, the Ontology becomes the primary interface for Operational AI.

Semantic Search via Embeddings

Traditional search relies on keyword matching. If you search for "engine failure," a keyword search might miss a report titled "Propulsion System Malfunction."

Semantic Search uses AI models to transform text into Embeddings—numerical vectors in N-dimensional space.

  • Vector Proximity: The system calculates the distance (often using Cosine Similarity) between the search query vector and the document vectors.
  • Ontology Integration: These vectors are stored as properties on Objects. This allows a user to search for "Safety Issues" and find "Aircraft" objects linked to "Maintenance Logs" that are semantically related to safety, even if the word "safety" is never used.

[AI_DEMO: A 2D projection of a high-dimensional vector space. Users can click on dots representing "Maintenance Logs." Logs with similar semantic meanings (e.g., "brake wear" and "stopping distance issues") cluster together, while unrelated logs (e.g., "catering schedule") are far apart.]

Models in the Ontology

In Palantir, AI/ML models are not isolated artifacts. They are integrated into the Ontology as Model Types. This allows a model to:

  1. Consume Objects: A "Churn Prediction Model" takes a Customer object as input.
  2. Produce Properties: The model outputs a Churn Probability property directly onto the Customer object.
  3. Trigger Actions: If the probability exceeds 0.8, the Ontology can trigger an Action to "Send Retention Email."

Ontology Design: Best Practices and Anti-Patterns

Building an Ontology is an exercise in domain modeling. Poor design leads to "System Silos" or "The Kitchen Sink" (where too much irrelevant data is included).

Best Practice: Model Reality, Not Systems

A common mistake is to mirror the structure of the source database (e.g., creating an object called SAP_T501_Table). Instead, the Ontology should reflect the business reality. If the business talks about "Purchase Orders," create a Purchase Order object, even if that data comes from five different tables in SAP.

Anti-Pattern: The God Object

The God Object is an Object Type that has hundreds of properties and dozens of links, attempting to represent everything about a business unit. This leads to performance degradation and cognitive overload for users.

  • Solution: Use Interfaces to break the God Object into smaller, functional components that can be composed as needed.

Anti-Pattern: System Silos

Creating separate "Finance Employee" and "HR Employee" objects because the data comes from different systems.

  • Solution: Use Data Integration to merge these into a single Employee object, using the Ontology to resolve conflicts and provide a unified view.

Summary of Foundations

The Palantir Ontology is the bridge between the "what" (data) and the "so what" (action). By combining semantic clarity with kinetic capability, it transforms a static data lake into a living digital twin.

[AI_FLASHCARDS:

  • Object Type: A digital representation of a real-world entity.
  • Link Type: A relationship between two Object Types.
  • Action Type: A predefined, validated way to change data in the Ontology.
  • Function: TypeScript code used to perform complex logic on Ontology objects.
  • Digital Twin: A virtual model of a physical process or object.
  • Semantic Search: Searching by meaning/context using vector embeddings.
  • Interface: A reusable template that defines properties and links for multiple Object Types.
  • Writeback: The process of sending changes made in the Ontology back to source systems.]

[AI_QUIZ:

  1. What is the primary difference between a Link Type and a standard Database Join?
    • A) Joins are faster, Links are slower.
    • B) Links are semantic, first-class entities that allow for object-based traversal without writing code. (Correct)
    • C) Links can only be 1:1, while Joins can be M:M.
  2. Why are Action Types considered "Kinetic"?
    • A) Because they move data between servers.
    • B) Because they represent the "verbs" or movements/changes within the organization. (Correct)
    • C) Because they are written in C++.
  3. Which anti-pattern is characterized by an Object Type having too many unrelated properties?
    • A) System Silo
    • B) The Kitchen Sink
    • C) The God Object (Correct)
  4. How does Semantic Search handle a query for "Rainy Weather" if the data only contains the word "Precipitation"?
    • A) It fails because the keywords don't match.
    • B) It succeeds because both terms are close to each other in N-dimensional vector space. (Correct)
    • C) It requires a manual synonym map to be created by an admin.]

[AI_STUDY_GUIDE: Foundations of the Palantir Ontology

  • Core Goal: Create an operational layer that functions as a digital twin.
  • Key Components:
    • Semantic: Objects (Nouns), Properties (Attributes), Links (Relationships).
    • Kinetic: Actions (State changes), Functions (Logic/Calculations).
  • Advanced Concepts:
    • Interfaces: Enable polymorphism and scale.
    • Semantic Search: Use embeddings for context-aware discovery.
    • Operational AI: Integrate models directly into the object graph.
  • Design Philosophy: Always model the real-world business process, not the underlying IT systems. Avoid "God Objects" and "System Silos" by using interfaces and intentional curation.]

Ontology Design and Best Practices

Key concepts: Model reality, not systems · Intentional curation · System Silos · The God Object · Department Silos

Best practices for modeling real-world entities and avoiding common architectural anti-patterns.

Ontology Design and Best Practices

In the landscape of modern enterprise data architecture, the Ontology serves as the critical operational layer that bridges the gap between raw digital assets and real-world decision-making. Unlike a traditional data warehouse, which focuses on the storage and retrieval of records, an Ontology is a digital twin of the organization. It is a rich, semantic representation of the entities, relationships, and processes that define a business.

Building an effective Ontology is not merely a technical exercise in data modeling; it is a strategic endeavor to create a shared language for the entire enterprise. When designed correctly, the Ontology enables connectivity at scale, allowing disparate data sources to converge into a unified, intuitive interface that empowers both human operators and artificial intelligence.

AI_SVGI_SVG## The Core Philosophy: Model Reality, Not Systems

The most fundamental principle of ontology design is to model reality, not systems. In many legacy environments, data structures are dictated by the underlying software that generated them—SAP, Salesforce, or custom SQL databases. This leads to a fragmented landscape where the same real-world entity (e.g., a "Customer") is represented in dozens of different, incompatible ways.

The Shift from System-Centric to Object-Centric

In a system-centric model, a user must understand the schema of the source database to find information. In an object-centric model (the Ontology), the user simply interacts with the "Aircraft" or "Employee" object.

Definition: The Digital Twin A digital twin is a virtual representation of a physical object or system. In the context of the Palantir Ontology, it is the mapping of integrated data (datasets, models, and virtual tables) to their real-world counterparts, including their properties, their relationships (links), and the actions that can be performed upon them.

Feature System-Centric (Legacy) Object-Centric (Ontology)
Primary Unit Table / Row Object / Entity
Language SQL / Technical Schema Business / Domain Language
Logic Location Application Code / ETL Actions & Functions (Kinetic Layer)
User Access Data Scientists / Engineers Operational Users / Analysts / AI
Integration Point-to-Point Unified Semantic Layer

Semantic Elements: The Building Blocks of Meaning

An Ontology is composed of Semantic Elements that define what things are and how they relate.

1. Object Types

Object types are the "nouns" of the Ontology. They represent discrete entities such as Equipment, Purchase Order, or Patient. Each object type is backed by a data source but abstracts away the complexity of that source.

2. Properties

Properties are the attributes of an object. For a Vehicle object, properties might include VIN, Color, and Last Maintenance Date. A key best practice here is Shared Properties, which allow different object types to use the same property definition (e.g., ID or Location), ensuring consistency across the graph.

3. Link Types

Links are the "connective tissue" of the Ontology. They define how objects relate to one another. Links can be:

  • One-to-One: A Pilot has one License.
  • One-to-Many: A Flight has many Passengers.
  • Many-to-Many: Parts are used in many Assemblies, and Assemblies contain many Parts.

Kinetic Elements: The Verbs of the Enterprise

A static map of data is useful for analysis, but an operational Ontology must be kinetic. Kinetic elements allow the Ontology to "do" things, transforming it from a passive reference into an active orchestration layer.

Action Types

Action types define how users and systems can modify the Ontology. They encapsulate business logic, validation rules, and permissions. Instead of writing a SQL UPDATE statement, a user triggers a "Reschedule Flight" action. This ensures that every change to the data is compliant with business rules and is captured in the Data Writeback layer.

Functions

Functions are units of logic (often written in TypeScript) that can be executed on the Ontology. They power dynamic properties, complex aggregations, and model inferences. For example, a function might calculate the "Risk Score" of a Supply Chain node based on real-time weather data and historical delay patterns.

// Example: A Kinetic Function to calculate Remaining Useful Life (RUL)
// This function sits in the Ontology layer, making the logic reusable across apps.

import { Function, Integer } from "@foundry/functions-api";
import { Equipment } from "@foundry/ontology-api";

export class EquipmentLogic {
    @Function()
    public calculateRUL(unit: Equipment): Integer {
        const currentHours = unit.operatingHours ?? 0;
        const maxHours = unit.designLifeHours ?? 10000;
        const degradationFactor = unit.environmentFactor ?? 1.0;
        
        const remaining = (maxHours - currentHours) / degradationFactor;
        return Math.max(0, Math.floor(remaining));
    }
}

Common Anti-Patterns in Ontology Design

Designing an Ontology is an iterative process, but certain "traps" can lead to architectural debt and user confusion.

1. System Silos

The Problem: The Ontology mirrors the source system exactly. If the source has a table called TMP_USER_V2_FINAL, the Ontology has an object type called Tmp User V2 Final. The Consequence: Users must still be experts in the source system to use the data. The "Digital Twin" fails because it is a twin of the database, not the business. The Fix: Use the Ontology Management Application (OMA) to rename, curate, and merge fields into business-friendly terms.

2. The God Object (The Kitchen Sink)

The Problem: Attempting to create a single object type that represents everything. For example, a Transaction object that contains 500 properties covering everything from shipping logistics to tax compliance to customer sentiment. The Consequence: Massive performance degradation, "property bloat," and extreme cognitive load for users. The Fix: Use Object Polymorphism and Interfaces. Break the God Object into smaller, specialized objects linked together, or use Interfaces to provide different "views" of the same entity for different departments.

3. Department Silos

The Problem: Designing the Ontology based on how a specific team (e.g., Marketing) views the data, ignoring how other teams (e.g., Finance) use the same entities. The Consequence: Duplicate objects (e.g., Marketing_Customer and Finance_Account) that never sync, leading to "multiple versions of the truth." The Fix: Implement Intentional Curation. Establish a cross-functional governance body to define core objects that serve the entire organization.

Anti-Pattern Symptom Solution
System Silos Object names like STG_SAP_MARA Map to real-world names like Material
The God Object 200+ properties on one object Use Links and Interfaces
Department Silos Client vs Customer vs Account Standardize via Shared Properties
The Kitchen Sink Every column in the raw data is a property Intentional Curation: Only include what is needed

Advanced Concept: Semantic Search and Embeddings

Modern Ontologies are increasingly used to power Ontology-Augmented Generation (OAG) and AI workflows. This requires moving beyond simple keyword matching to Semantic Search.

Semantic search uses AI models to transform text into numerical vectors (embeddings) in an N-dimensional space. In this space, the distance between vectors represents the similarity in meaning.

The Process of Semantic Integration:

  1. Chunking: Large documents (PDFs, manuals) are broken into semantically distinct "chunks."
  2. Embedding: Each chunk is passed through an embedding model (like OpenAI Ada or MSMARCO) to create a vector.
  3. Indexing: These vectors are stored as properties on Ontology objects.
  4. Retrieval: When a user asks a question, the query is embedded, and the system finds the "nearest neighbors" in the vector space.

AI_DEMOI_DEMO### Document Processing with Pix2Struct For visually-situated language (like diagrams, forms, or screenshots), traditional OCR often fails. Models like Pix2Struct use a pixel-level Vision Transformer (ViT) to parse images directly into structured text or HTML. Integrating these models into the Ontology allows the digital twin to "see" and interpret complex visual data, such as engineering schematics or insurance claim photos.

Best Practices for Scalable Design

1. Intentional Curation

Do not ingest data just because it exists. Every object, property, and link in the Ontology should have a "reason for being"—a specific decision it supports or a workflow it enables. This reduces noise and maintenance overhead.

2. Interfaces for Abstraction

Interfaces allow you to define a set of properties and behaviors that multiple object types can implement. For example, an Asset interface might require a Location and Status property. Both Truck and Generator objects can implement Asset. Applications can then be built to work with any Asset, regardless of its specific type.

3. Action Types vs. Data Pipelines

A common mistake is using data pipelines (ETL) for operational updates.

  • Pipelines are for high-volume, batch processing of raw data into the Ontology.
  • Action Types are for user-driven, real-time changes. Using Actions ensures that the "human-in-the-loop" decisions are captured immediately and can trigger downstream logic or notifications.

4. Observability and Data Lineage

The Ontology Management Application (OMA) provides built-in observability. Always monitor the "health" of your object types. If a source dataset fails to update, the Ontology should reflect that the data is "stale." High-quality ontologies maintain clear Data Lineage, allowing any user to trace a property back to its source system of record.

Operationalizing AI/ML via the Ontology

The Ontology is the "last mile" for AI/ML. Most data science projects fail because the models remain isolated from the business process. By integrating models directly into the Ontology:

  • Live Inference: Models can run on-demand when a user views an object.
  • Decision Capture: When a model makes a recommendation (e.g., "Predictive Maintenance Required"), the user's response to that recommendation is written back into the Ontology, creating a feedback loop for model retraining.
  • Interpretability: Models operate on business objects (e.g., Churn Risk for a Customer), making the outputs immediately understandable to non-technical stakeholders.

Key Insight: The Feedback Loop The true power of an Ontology is not just in organizing data, but in capturing the decisions made based on that data. This "Decision Capture" creates a new, highly valuable data asset that records how the organization actually functions, which is the ultimate training data for future AI.

Ontology Design and Best Practices - Ontology building - diagram 1
Ontology Design and Best Practices - Ontology building - diagram 1

The Ontology Manager (OMA)

Key concepts: Ontology Management Application (OMA) · Review Edits Dialog · Ontology Metrics · Ontology Cleanup · JSON Schema Export/Import

A guide to the central application for building, maintaining, and monitoring the Ontology.

The Ontology Manager (OMA)

In the architecture of a modern data-driven enterprise, the transition from raw, siloed datasets to an integrated, actionable "Digital Twin" is facilitated by the Ontology. If the Ontology is the "brain" of the organization—storing the semantic definitions of every asset, person, and process—then the Ontology Management Application (OMA) is its primary interface and control plane.

The OMA is not merely a schema editor; it is a sophisticated development environment designed to manage the lifecycle of an organization's operational layer. It sits at the intersection of data engineering and application development, providing the tools necessary to map physical data assets (datasets, streams, and models) to logical entities (Objects, Properties, and Links). This section explores the mechanics of the OMA, its transactional integrity models, and the advanced workflows required to maintain a high-performance Ontology at scale.

AI_SVGI_SVG## The Ontology as an Operational Layer

To understand the OMA, one must first grasp the concept of the Operational Layer. Traditional data architectures often suffer from a "semantic gap" between the data warehouse (optimized for storage and SQL queries) and the end-user application (optimized for business logic). The Palantir Ontology bridges this gap by providing a rich, object-oriented representation of the organization.

Definition: The Ontology is a formal, semantic representation of an organization's entities (Object Types), their attributes (Properties), their relationships (Link Types), and the logic that governs them (Action Types and Functions).

In this paradigm, the OMA serves as the "compiler" and "orchestrator" for this layer. It ensures that every change—whether adding a new Tail Number property to an Aircraft object or defining a complex Dispatch action—is validated against the underlying data lineage and the downstream application requirements.


The Architecture of the Ontology Management Application (OMA)

The OMA is structured to support both exploratory design and rigorous production management. It provides a multi-dimensional view of the organization’s data health, connectivity, and usage.

The Discover View and Object Type Management

The OMA’s Discover View acts as the entry point for architects. It provides a high-level visualization of the entity-relationship diagram (ERD) of the organization. Unlike static documentation, this view is "live," reflecting the actual state of data synchronization and indexing.

Component Functionality Primary User
Object Type View Defines the mapping between a backing dataset and a semantic object. Data Architect
Action Type View Defines the "Kinetic" elements—how users can modify the Ontology through validated logic. App Developer
Link Type View Establishes the cardinality (1:1, 1:N, M:N) and foreign-key relationships between objects. Data Architect
Function View Manages TypeScript-based logic for computed properties and complex aggregations. Software Engineer

The Transactional Model: Review Edits Dialog

One of the most critical features of the OMA is its transactional integrity. Changes made within the OMA are not applied immediately to the production environment. Instead, they are stored in a "Draft" state, allowing for complex, multi-object refactoring without risking system stability.

The Review Edits Dialog serves as the final gatekeeper. It performs a series of automated checks:

  1. Schema Validation: Ensures that property types (e.g., Integer, String, Geopoint) match the schema of the backing datasets.
  2. Dependency Analysis: Identifies if removing a property will break downstream Workshop applications or Quiver analyses.
  3. Permission Auditing: Verifies that the user has the requisite Ontology Lead or Editor roles to modify specific namespaces.

AI_DEMOI_DEMO--

Ontology Metrics: The Observability of Meaning

A common failure in enterprise data management is the "Build it and they will come" fallacy. To prevent this, the OMA provides Ontology Metrics, a robust observability suite that tracks the "pulse" of the operational layer.

Read and Write Dynamics

The OMA tracks metrics over a rolling 30-day window, providing a quantitative basis for architectural decisions.

  • Reads: Measures how often an object type or link is queried by end-user applications (e.g., an "Active Alerts" dashboard). High read volume on a specific object suggests it is a "Core Entity."
  • Writes: Measures the frequency of Action Types being executed. This represents the "Kinetic" energy of the Ontology—how much the organization is actually changing its state through the platform.

Metric-Driven Optimization

By analyzing the ratio of Reads to Writes, architects can identify "Dead Objects" (high storage cost, low usage) or "Bottleneck Objects" (high write contention).

Metric Category Metric Name Definition Strategic Value
Usage Active Users Unique users interacting with the object type. Identifies stakeholder alignment.
Performance Query Latency Time taken to resolve object-set filters. Triggers indexing optimization.
Volume Instance Count Total number of instantiated objects (rows). Manages storage and compute costs.
Connectivity Link Density Average number of links per object instance. Measures the "connectedness" of the twin.

Advanced Editing: JSON Schema Export and Import

For senior engineers, the OMA provides a "headless" or programmatic interface via JSON Schema Export/Import. While the GUI is excellent for visualization, bulk operations—such as migrating an entire Ontology from a Development to a Production environment or renaming fifty properties—are best handled through the schema.

The Anatomy of an Object Definition

The Ontology is serialized as a JSON object. This allows for version control (e.g., storing the schema in a Git repository) and CI/CD integration.

{
  "objectTypes": {
    "aircraft": {
      "apiName": "aircraft",
      "displayName": "Aircraft",
      "primaryKey": "tailNumber",
      "properties": {
        "tailNumber": { "type": "string", "indexed": true },
        "lastMaintenanceDate": { "type": "date" },
        "currentLocation": { "type": "geopoint" }
      },
      "backingDatasource": {
        "type": "dataset",
        "rid": "ri.foundry.main.dataset.12345"
      }
    }
  },
  "linkTypes": {
    "aircraftToFlight": {
      "apiName": "aircraft_flights",
      "cardinality": "ONE_TO_MANY",
      "links": [
        { "objectType": "aircraft", "property": "tailNumber" },
        { "objectType": "flight", "property": "aircraftId" }
      ]
    }
  }
}

Programmatic Workflows

By exporting the JSON, an engineer can use standard text-processing tools (Python, jq) to perform mass updates. For example, if a company rebrands and needs to change all apiNames from legacy_system_* to modern_core_*, a simple regex script on the JSON schema is significantly faster and less error-prone than manual GUI edits.


Ontology Cleanup and Technical Debt Management

As an organization grows, its Ontology tends toward entropy. New projects add "temporary" objects that eventually become permanent technical debt. The Ontology Cleanup Tool is the OMA's primary mechanism for maintaining a "Lean Ontology."

Identifying Candidates for Deprecation

The OMA uses a heuristic approach to suggest objects for cleanup:

  1. Orphaned Objects: Object types with no active links to other entities.
  2. Stale Data: Objects whose backing datasets haven't been updated in over 90 days.
  3. Low-Utility Objects: Objects with zero "Reads" in the last 30 days.

The "God Object" Anti-Pattern

A frequent pitfall in Ontology design is the creation of a God Object—a single object type (e.g., Global_Asset) that attempts to capture every possible property for every different type of asset. This leads to sparse data, poor performance, and confusing user interfaces. The OMA's Interface and Polymorphism features allow architects to break these down into specialized object types that share a common contract, ensuring both specificity and interoperability.

Key Insight: A healthy Ontology is not the one with the most objects, but the one with the highest semantic density—where every object and link serves a specific, documented operational purpose.


Semantic Search and Embedding Integration

Modern Ontologies are increasingly "AI-aware." The OMA now includes configurations for Semantic Search, allowing users to search for objects based on meaning rather than exact keyword matches.

Mechanics of Vector Integration

In the OMA, a property can be designated as "Searchable via Embeddings." When this is enabled:

  1. The OMA triggers a process to pass the text property (e.g., Maintenance_Notes) through an Embedding Model (like OpenAI's Ada or a local LLM).
  2. The resulting N-dimensional vector is stored in a specialized index.
  3. At query time, the OMA facilitates a "Nearest Neighbor" search in vector space.

Worked Example: Semantic Proximity

Imagine an object type Work_Order with a property Description.

  • Keyword Search: Searching for "Engine" only returns orders containing that exact word.
  • Semantic Search: Searching for "Propulsion issues" returns orders mentioning "turbine failure," "thrust loss," or "fan blade damage," even if the word "propulsion" is absent.

The OMA manages the mapping of these embedding models to the properties, ensuring that as the underlying data updates, the vector index is re-calculated automatically.


Common Pitfalls in OMA Management

Even with a powerful tool like the OMA, architectural mistakes can occur. Senior practitioners should watch for the following:

Pitfall Description Consequence Mitigation
Modeling the Source, Not Reality Creating object types that mirror the messy structure of the source SQL database. Users cannot understand the data; logic is brittle. Model based on real-world entities (e.g., "Pump" not "TBL_PMP_V2").
Ignoring Link Cardinality Setting all links to M:N (Many-to-Many) because it's "easier." Massive performance degradation during joins. Use 1:N whenever possible; enforce referential integrity.
Over-Indexing Marking every property as "Searchable" or "Sortable." Increased storage costs and slower write speeds. Only index properties used in filters or search bars.
Breaking Changes Deleting a property that is a "Primary Key" for a downstream Action. Production applications will crash immediately upon save. Use the "Review Edits" dialog to check for "Usage References" before deleting.

Summary: The OMA as the Heart of the Digital Twin

The Ontology Management Application is the bridge between the static world of data engineering and the dynamic world of organizational operations. By providing a transactional, observable, and programmatic environment, it allows organizations to build a "Digital Twin" that is not just a visualization, but a living asset.

As organizations move toward Ontology-Augmented Generation (OAG) and LLM-driven workflows, the precision of the OMA becomes even more critical. An LLM is only as effective as the context it is given; a well-managed Ontology, curated via the OMA, provides the "Ground Truth" that prevents AI hallucinations and ensures that automated decisions are based on the actual state of the enterprise.

AI_STUDY_GUIDEI_STUDY_GUIDE### Further Reading and Formalisms For those interested in the mathematical underpinnings of the OMA's mapping logic, consider the study of Category Theory as applied to database schemas. The OMA essentially performs a functorial mapping from the category of Datasets (where morphisms are joins/transforms) to the category of Objects (where morphisms are semantic Links). Ensuring the commutativity of these mappings is the fundamental task of the Ontology Architect.

The Ontology Manager (OMA) - Ontology building - diagram 1
The Ontology Manager (OMA) - Ontology building - diagram 1

Permissions and Governance

Key concepts: Project-based permissions · Resource inheritance · Ontology Roles · Metadata vs. Data Decoupling · Migration Assistant

Understanding the transition to project-based permissions and managing ontology roles.

Permissions and Governance

The Palantir Ontology is more than a simple schema; it is the operational layer of an organization. Because it sits between raw data assets (datasets, virtual tables, models) and end-user applications (Workshop, Quiver, Slate), its security model must be both robust enough to protect sensitive information and flexible enough to enable cross-functional collaboration. Governance in the Ontology is the practice of defining who can see the "digital twin" of the organization, who can modify its structure, and who can trigger "kinetic" changes (Actions) that write back to source systems.

AI_SVGI_SVGThe Infographic should depict the "Three-Layer Security Architecture": 1. The Data Layer (Datasets/Foundry Files), 2. The Semantic Layer (Object Types/Link Types), and 3. The Application Layer (Action Types/Workshop Modules). It should show how permissions flow from Projects down to individual resources via Resource Inheritance.*

Project-Based Permissions

In the early iterations of the platform, permissions were often managed at the individual resource level or derived directly from the underlying data sources. As organizations scaled to tens of thousands of object types, this became unmanageable. Modern Palantir governance utilizes a Project-based permissions model.

What it is

Project-based permissions centralize the management of security within Compass Projects. A Project acts as a security boundary—a container for datasets, object types, action types, and logic. Instead of assigning a user "Viewer" access to a specific "Flight" object, the user is granted a role on the "Aviation Operations" Project, which then grants them access to all ontology elements contained within.

Why it matters

Managing permissions at scale requires a shift from micro-management to macro-governance.

  1. Scalability: Administrators manage hundreds of projects rather than millions of individual files.
  2. Consistency: Ensures that if a user can see the data, they can also see the semantic representation of that data.
  3. Auditability: Simplifies the process of answering "Who has access to our customer data?" by looking at Project memberships.

How it works: Resource Inheritance

The core mechanism of this model is Resource Inheritance. When an Object Type or Action Type is created within a Project, it automatically inherits the security settings of that Project.

Definition: Resource Inheritance The principle where a child resource (e.g., an Object Type) automatically adopts the Access Control Lists (ACLs) and security markings of its parent container (the Project). This creates a "secure by default" environment.

Feature Legacy (Datasource-Derived) Modern (Project-Based)
Primary Boundary The underlying Dataset The Compass Project
Management Unit Individual Files Logical Groupings (Projects)
Inheritance Direction Bottom-Up (Data -> Ontology) Top-Down (Project -> Resources)
Complexity High (n resources to manage) Low (m projects to manage)
Flexibility Rigid; tied to data storage High; tied to business use-case

Ontology Roles and Granular Access

While Projects provide the boundary, Ontology Roles provide the granularity. Not every user in a project should have the same level of control over the Ontology.

What they are

Ontology Roles are a set of predefined permissions that dictate what a user can do within the Ontology Management Application (OMA) and within ontology-aware applications. These roles are layered on top of standard Project roles (Viewer, Editor, Owner).

Key Roles and Capabilities

The system distinguishes between those who define the world (Architects), those who manage the data (Data Engineers), and those who operate within it (End Users).

Role Primary Function Key Capabilities
Ontology Viewer Read-only access View object definitions, browse the OMA, use objects in Quiver/Workshop.
Ontology Editor Structural modification Create/Edit Object Types, Link Types, and Action Types. Map properties to datasets.
Action Executor Operational use Permission to trigger specific Action Types that modify data or trigger external systems.
Metadata Manager Governance Edit descriptions, tags, and display metadata without changing the underlying schema.

Common Pitfall: The "Kitchen Sink" Project

A common mistake is placing all Ontology elements into a single "Global Ontology" Project. While this seems simple, it breaks the principle of Least Privilege. If a user needs access to "Public Transport" data, they should not necessarily inherit "Payroll" object types just because they are in the same project. Best practice suggests grouping objects by functional domain (e.g., "Supply Chain," "Human Resources").


Metadata vs. Data Decoupling

One of the most sophisticated aspects of Palantir governance is the decoupling of metadata from data. This allows for a "Discovery" workflow where users can understand the organization's data landscape without necessarily having access to the sensitive records themselves.

The Concept

  • Metadata: The definition of the Object Type (e.g., "An 'Employee' object has an 'ID', a 'Name', and a 'Salary' property").
  • Data: The actual instances of that object (e.g., "Employee #123, John Doe, $100,000").

Why it matters

Decoupling enables Interpretability and Economies of Scale. A data scientist can browse the Ontology to see what data is available (Metadata) to request access to the specific datasets they need (Data). This prevents the "Silo" effect where users don't even know what data exists because they don't have permission to see it.

Implementation Mechanics

In the OMA, an administrator can grant a user the ability to "Discover" an object type. This allows the user to see the object in the Object Explorer and see its properties, but when they attempt to run a query or view a specific record, the system checks the underlying Datasource Permissions. If the user lacks access to the backing dataset, the records remain hidden.

Theorem of Ontology Access Let $A_{total}$ be the total access a user has to an object. $A_{total} = P(Metadata) \cap P(Data)$ Where $P(Metadata)$ is the permission to see the definition in the Project, and $P(Data)$ is the permission to see the records in the underlying dataset/marking.


Governance of Kinetic Elements: Action Types and Functions

The Ontology is not just a static "Digital Twin"; it is an operational layer that includes Kinetic Elements: Action Types and Functions. Governing these is critical because they represent the ability to change the state of the organization.

Action Type Permissions

Action Types allow users to modify properties or create links. Because these can have real-world consequences (e.g., "Approve Loan," "Dispatch Technician"), their governance is multi-layered:

  1. Discovery: Who can see that this action exists?
  2. Execution: Who has the right to trigger the action?
  3. Validation: Logic-based checks (Functions) that ensure the action is "legal" according to business rules.

Code Example: Action Validation Logic

The following logic demonstrates how a Function can be used as a "Governance Guard" within an Action Type.

import { OntoAction, Employee, LoanApplication } from "@foundry/ontology-api";

export class LoanActions {
    /**
     * @function approveLoan
     * Validates if the current user has the authority to approve a specific loan amount.
     * This acts as a governance layer above simple project permissions.
     */
    @OntoAction()
    public approveLoan(loan: LoanApplication, approver: Employee): void {
        // Governance Check: Approver must have a 'Seniority' level > 5 for loans over $50k
        if (loan.amount > 50000 && approver.seniorityLevel <= 5) {
            throw new Error("Governance Violation: Insufficient seniority for this loan amount.");
        }

        // Logic to update the object
        loan.status = "Approved";
        loan.approvalDate = new Date();
    }
}

Why this matters

By embedding governance logic directly into Functions, organizations move from "Static Permissions" (Yes/No) to "Contextual Governance" (Yes, if conditions A, B, and C are met). This is essential for Decision Capture and maintaining a high-fidelity audit trail.


The Migration Assistant

As Palantir evolved its security model, many organizations found themselves with "Legacy" ontologies where permissions were fragmented. The Migration Assistant is the specialized tool designed to transition these resources to the modern Project-based model.

The Migration Workflow

  1. Analysis: The tool scans the existing Ontology to identify "Datasource-derived" objects.
  2. Mapping: The administrator maps these objects to their new parent Projects.
  3. Validation: The system checks for "Permission Gaps"—cases where a user might lose access or gain unintended access during the move.
  4. Execution: The tool updates the internal pointers (the "Resource Identifiers") to point to the Project's ACLs.

Pitfalls in Migration

  • Circular Dependencies: Object A in Project 1 links to Object B in Project 2, but Project 2 requires a reference back to Project 1. This can create complex "Permission Loops."
  • Broken Applications: If an application (e.g., a Workshop module) is not in the same project as the objects it uses, and "Reference" permissions aren't set, the app will fail to load.

Best Practices for Ontology Governance

Effective governance requires a balance between strict control and user enablement. The following table summarizes best practices derived from large-scale implementations.

Practice Description Benefit
Intentional Curation Only promote high-quality, governed datasets to the Ontology. Prevents the "Kitchen Sink" anti-pattern and maintains trust.
Interface Abstraction Use Interfaces to define common properties across different object types. Allows for Object Polymorphism and consistent security application.
Separation of Concerns Keep Logic (Functions) in separate projects from Data (Datasets). Enables developers to iterate on logic without needing access to PII.
Action-First Writeback Always use Action Types for data modification rather than direct dataset edits. Ensures all changes are validated and captured in the Digital Twin audit log.

AI_DEMOI_DEMOThe Demo would be a "Permission Simulator." A user can select a "User Role" (e.g., Analyst), a "Project Role" (e.g., Viewer), and a "Security Marking" (e.g., Restricted). The simulator would then highlight which Objects and Actions in a sample "Supply Chain" Ontology are visible, hidden, or "Discoverable-only." This helps visualize the intersection of Metadata and Data permissions.*

Conclusion: The Operational Layer as a Trust Anchor

Governance in the Palantir Ontology is what transforms a "Data Lake" into an "Operational Platform." By decoupling metadata from data, utilizing project-based inheritance, and enforcing logic through kinetic elements, organizations can create a Digital Twin that is both highly accessible and rigorously secure. This "Trust Anchor" allows non-technical users to make data-driven decisions with the confidence that the underlying security and business logic are being enforced automatically.

AI_FLASHCARDSI_FLASHCARDS Project-based Permissions: A security model where ontology resources inherit ACLs from their parent Compass Project.

  • Resource Inheritance: The mechanism by which child objects automatically adopt the security posture of their parent container.
  • Metadata vs. Data Decoupling: The separation of an object's definition (schema) from its actual records (content), allowing for discovery without full access.
  • Ontology Roles: Specific permissions (Viewer, Editor, Action Executor) that define user capabilities within the Ontology layer.
  • Kinetic Elements: Components of the Ontology (Actions and Functions) that enable state changes and operational logic.
  • Migration Assistant: A tool used to move legacy ontology resources into the modern project-based security framework.
  • Object Polymorphism: The ability for different object types to be treated as a single type via a shared Interface, simplifying governance.

AI_QUIZI_QUIZ. True or False: In the modern Palantir security model, if you have access to a dataset, you automatically have the "Ontology Editor" role for any object type mapped to that dataset. (Answer: False. Roles are managed at the Project level, and Editor status is a specific assignment.) 2. Which concept allows a user to see that a "Credit Card" object exists without seeing actual credit card numbers?

  • A) Resource Inheritance
  • B) Metadata vs. Data Decoupling
  • C) Action Validation
  • D) The Migration Assistant (Answer: B)
  1. What is the primary risk of the "God Object" anti-pattern in governance?
    • A) It makes the data load too slowly.
    • B) It forces too many users to have access to a single, over-privileged object, violating Least Privilege.
    • C) It prevents the use of TypeScript functions. (Answer: B)
  2. How does an "Action Type" differ from a standard data pipeline in terms of governance?
    • A) Actions are only for read-only data.
    • B) Actions provide real-time validation and "Decision Capture" within the operational UI.
    • C) Actions do not require permissions. (Answer: B)

AI_STUDY_GUIDEI_STUDY_GUIDE*Key Themes to Master:**

  • The Shift in Security Philosophy: Understand the transition from datasource-centric security to project-centric security. Why was this necessary for the "Operational Layer"?
  • The "Digital Twin" Concept: How do semantic elements (Objects, Links) and kinetic elements (Actions) work together to represent a business process?
  • Granularity vs. Usability: How do Ontology Roles and Discovery permissions balance the need for security with the need for data democratization?
  • Implementation Path: If you were tasked with moving a legacy organization to the modern Ontology, what steps would you take using the Migration Assistant?
  • Logic-Based Governance: Be prepared to explain how Functions serve as a "Validation Layer" for Actions, moving beyond simple binary permissions.
Permissions and Governance - Ontology building - diagram 1
Permissions and Governance - Ontology building - diagram 1

Ontology-Aware Applications and Logic

Key concepts: Workshop · Quiver · Object Explorer · AI/ML Operationalization · Derived properties · Ontology SDK

How to use the Ontology to power analytical tools, AI models, and dynamic logic.

Ontology-Aware Applications and Logic

The Palantir Ontology serves as the central nervous system of the modern enterprise, acting as a high-fidelity digital twin that maps disparate data assets—datasets, virtual tables, and streaming inputs—into a unified semantic framework. However, an ontology is not merely a static map; it is an operational layer. To derive value from this layer, organizations utilize Ontology-Aware Applications, which are software environments designed to "speak" the language of the Ontology natively.

By abstracting the underlying complexity of SQL joins and distributed storage, these applications allow users to interact with real-world entities—such as Aircraft, Work Orders, or Patients—and their relationships. This section explores the mechanics of these applications, the logic that governs them, and the advanced operationalization of AI/ML models within this framework.

AI_SVGI_SVGThe Ontology Pipeline: From Raw Data Integration to Semantic Modeling, Kinetic Logic (Actions/Functions), and finally to Ontology-Aware Consumption in Workshop, Quiver, and the SDK.*


The Ontology as an Operational Layer

Before diving into specific applications, we must define the Ontology in its operational context. It is the bridge between the "Data Lake" (a passive repository) and the "Operating System" (an active decision-making environment).

Definition: The Operational Ontology The Ontology is a formal representation of an organization's entities (Object Types), their attributes (Properties), their interconnections (Link Types), and the valid transformations they can undergo (Action Types). It provides a semantic contract that ensures every application built on top of it shares a common understanding of the business logic.

Semantic vs. Kinetic Elements

The Ontology is divided into two primary dimensions:

  1. Semantic Elements: The "nouns" and "adjectives" (Objects and Properties).
  2. Kinetic Elements: The "verbs" (Actions and Functions).
Element Type Component Description Operational Role
Semantic Object Type A template for a real-world entity (e.g., Employee). Defines data structure and permissions.
Semantic Property A discrete attribute of an object (e.g., Salary). Provides the basis for filtering and aggregation.
Semantic Link Type A relationship between objects (e.g., Works_At). Enables graph-based traversal and joins.
Kinetic Action Type A predefined logic for modifying data (e.g., Promote). Orchestrates write-backs and side effects.
Kinetic Function Logic-based computations (e.g., Calculate_Risk). Powers derived properties and complex validations.

Object Explorer: The Set-Theoretic Search Engine

Object Explorer is the primary interface for discovery within the Ontology. Unlike traditional Business Intelligence (BI) tools that query tables, Object Explorer operates on Object Sets.

How it Works: The Exploration Tree

Object Explorer uses a set-theoretic approach to data discovery. Every interaction—filtering by a property, following a link, or applying a search—is a transformation on a set.

  • Initial Set ($S_0$): All instances of an Object Type.
  • Filter Transformation ($f$): $S_1 = {x \in S_0 \mid P(x)}$, where $P$ is a predicate.
  • Pivot Transformation ($g$): $S_2 = {y \mid \exists x \in S_1, L(x, y)}$, where $L$ is a Link Type.

This allows users to perform "linked analysis." For example, a user can start with a set of Delayed Flights, pivot to the Aircraft performing those flights, and then pivot again to the Maintenance Logs for those specific aircraft.

Search Syntax and Indexing

Object Explorer supports advanced search capabilities, including Regular Expressions (Regex) and Semantic Search. For Regex to be performant at scale, properties must be specifically indexed in the Ontology Management Application (OMA). This indexing creates inverted indices that allow for sub-second retrieval across millions of objects.


Quiver: Advanced Analytical Orchestration

While Object Explorer is for discovery, Quiver is for deep analysis. It treats the Ontology as an Object-Graph, allowing for time-series analysis, regression, and complex transformations.

Key Mechanics: The Actionable Canvas

Quiver allows users to build analytical "cards" that are reactive. If the input Object Set changes, the entire downstream analysis updates. This is particularly powerful for Time-Series data. In the Ontology, a property can be a "Time Series" type, linking an object to a history of values (e.g., a Sensor object linked to Temperature readings over time).

Why it Matters

Quiver solves the "siloed analysis" problem. In traditional environments, an analyst might download a CSV, perform a regression in Excel, and then email a screenshot. In Quiver, the analysis is ontology-backed. If the underlying data for a Sensor is updated in the data pipeline, the regression model in Quiver reflects that change immediately.


Workshop: Building Operational Applications

Workshop is the premier environment for building high-fidelity, "low-code" applications. It is designed for Decision-Making Orchestration.

Mechanics of Workshop

Workshop applications are built around Events and Actions.

  1. State Management: Workshop maintains a "current state" of selected objects.
  2. Widgets: UI components (tables, maps, charts) that react to the state.
  3. Action Integration: The core of Workshop is the ability to trigger Action Types.

Action Types and Data Writeback

Action Types are the mechanism for "writing back" to the Ontology. When a user clicks a button in Workshop to "Approve a Loan," the Action Type:

  • Validates the user's permissions.
  • Checks logic constraints (e.g., "Is the loan amount < $1M?").
  • Updates the Status property of the Loan object.
  • Logs the decision for auditability (Decision Capture).

AI_DEMOI_DEMOInteractive Simulation: A Workshop module where selecting a 'Factory' object on a map filters a 'Production Line' table, and clicking 'Maintenance Required' triggers an Action Type that updates the Ontology and sends a notification to a downstream system.*


AI/ML Operationalization: Models in the Ontology

One of the most advanced features of the Palantir Ontology is the integration of Machine Learning models as "first-class citizens." This is known as AI/ML Operationalization.

The Digital Twin of the Model

In traditional architectures, models sit in a "black box" (e.g., a SageMaker endpoint). In the Ontology, a model is integrated as a Model Objective or a Function.

  • Live Inference: Applications like Workshop can call a model in real-time. When a user inputs data into a form, the Ontology can pass that data to a model and display a Risk Score immediately.
  • Model-as-a-Function: Models are wrapped in Functions on Objects (FoO). This allows the model to take an Object as an input, read its properties, and return a prediction.

Worked Example: Predictive Maintenance

Consider a Turbine object.

  1. Input: The Turbine object has properties like Inlet Temperature, Vibration Level, and Operating Hours.
  2. Logic: A TypeScript Function is defined that gathers these properties and sends them to a deployed Random Forest model.
  3. Output: The function returns a Probability of Failure.
  4. Application: A Workshop app displays this probability. If it exceeds 0.8, the app enables a "Schedule Inspection" button (an Action Type).

Derived Properties and Logic Functions

Derived Properties (currently in beta) represent a shift from "stored data" to "computed data."

What they are

A standard property is stored in a database column. A Derived Property is calculated at runtime using a Function.

Theorem: The Computational Trade-off Let $C_s$ be the cost of storage and $C_c$ be the cost of runtime computation. A property $P$ should be derived if $C_c < C_s + C_{sync}$, where $C_{sync}$ is the latency cost of keeping stored data updated via batch pipelines.

Implementation via Ontology SDK

The Ontology SDK allows developers to interact with these properties and functions in a type-safe environment. When the Ontology is updated, the SDK generates updated TypeScript or Java classes.

// Example: Using the Ontology SDK to calculate a derived "Urgency Score"
import { Objects } from "@foundry/ontology-api";

export class TicketLogic {
    @Function()
    public async calculateUrgency(ticket: Objects.SupportTicket): Promise<number> {
        const customer = await ticket.customer.get();
        const baseUrgency = ticket.priority === "High" ? 10 : 5;
        
        // Logic: VIP customers get a multiplier
        if (customer?.status === "VIP") {
            return baseUrgency * 2;
        }
        return baseUrgency;
    }
}

Semantic Search and Vector Embeddings

Modern Ontologies must handle unstructured data (PDFs, images, logs). This is achieved through Semantic Search.

The Mechanics of Embeddings

Semantic search does not look for word matches; it looks for conceptual proximity.

  1. Chunking: Large documents are broken into smaller, semantically coherent "chunks."
  2. Vectorization: Each chunk is passed through an embedding model (e.g., OpenAI Ada, MSMARCO) to produce a vector in N-dimensional space.
  3. Ontology Integration: These vectors are stored as properties on "Chunk" objects, linked to the parent "Document" object.
  4. Querying: When a user searches, their query is also vectorized. The system performs a Cosine Similarity search to find the closest vectors.

$$\text{similarity} = \cos(\theta) = \frac{\mathbf{A} \cdot \mathbf{B}}{|\mathbf{A}| |\mathbf{B}|}$$

Multimodal Models

Applications can now use models like Pix2Struct to parse screenshots or diagrams into the Ontology. This allows "visually-situated language" (e.g., a chart in a PDF) to be converted into structured Object properties.


Best Practices and Anti-Patterns

Building an Ontology-aware ecosystem requires disciplined design.

Practice Description Why it matters
Model Reality, Not Systems Define objects based on business entities, not source database tables. Prevents "System Silos" and makes the Ontology intuitive for users.
Intentional Curation Only add properties and objects that serve a specific operational use case. Avoids the "Kitchen Sink" anti-pattern which degrades performance.
Interfaces for Abstraction Use Object Interfaces to define common behaviors (e.g., all Vehicles have a GPS_Location). Enables polymorphic applications that work across different object types.
Action Types over Pipelines Use Actions for user-driven changes rather than waiting for a batch ETL. Enables "Live" operations and immediate feedback loops.

Common Pitfall: The God Object

A common mistake is creating a God Object—a single object type (like Project) that has hundreds of properties and dozens of links. This leads to massive performance degradation in Object Explorer and makes the UI overwhelming. Instead, use Composition: break the God Object into smaller, linked sub-objects (e.g., ProjectDetails, ProjectFinances, ProjectTeam).


Summary of Application Capabilities

Application Primary User Core Logic Unit Best For
Object Explorer Business User Object Sets Discovery, Filtering, Basic Reporting.
Quiver Analyst Logic Cards / Graphs Root Cause Analysis, Time-Series, Statistics.
Workshop App Builder Actions / Widgets Operational Workflows, Decision Capture.
Slate Developer SQL / JS / APIs Highly bespoke, pixel-perfect dashboards.
Ontology SDK Software Engineer TypeScript / Java Custom web apps, external integrations.

AI_FLASHCARDSI_FLASHCARDS Object Set: A collection of objects of a specific type, often defined by a series of filters or pivots.

  • Action Type: A curated logic block that defines how users can modify the Ontology, including validations and side effects.
  • Function on Objects (FoO): A snippet of code (usually TypeScript) that executes logic using Ontology objects as inputs.
  • Digital Twin: A virtual representation of a physical asset or process, updated with real-time data.
  • Semantic Search: Searching by meaning rather than keywords, powered by vector embeddings.
  • Derived Property: A property whose value is calculated dynamically at runtime rather than stored in a static table.
  • Polymorphism (Interfaces): The ability for different object types to be treated as the same type if they share a common interface.

AI_QUIZI_QUIZ. Scenario: You need to build a tool for flight controllers to reassign pilots to flights. Which application and kinetic element are most appropriate?

  • A) Quiver with a Regression Card.
  • B) Workshop with an Action Type.
  • C) Object Explorer with a Filter.
  • D) Slate with a SQL query.
  • Answer: B. Workshop provides the UI, and Action Types handle the logic of reassignment.
  1. Conceptual: Why is "Model-as-a-Function" superior to a standalone API for operational AI?

    • A) It is faster to compute.
    • B) It bypasses security protocols.
    • C) It allows the model to natively understand the relationships (links) between objects in the Ontology.
    • D) It eliminates the need for data integration.
    • Answer: C. By being ontology-aware, the model can traverse links to gather context.
  2. Technical: In Semantic Search, what is the purpose of "Chunking"?

    • A) To compress data for storage.
    • B) To fit text within the token limits of embedding models and improve retrieval relevance.
    • C) To encrypt sensitive information.
    • D) To convert text into SQL tables.
    • Answer: B.

AI_STUDY_GUIDEI_STUDY_GUIDE*Key Learning Objectives:**

  • Understand the distinction between Semantic (Objects/Links) and Kinetic (Actions/Functions) elements.
  • Master the Set-Theoretic approach used in Object Explorer for data discovery.
  • Explain how AI/ML models are operationalized via Functions on Objects to provide live inference.
  • Differentiate between Workshop (low-code app building) and Quiver (analytical exploration).
  • Recognize the God Object anti-pattern and how to avoid it through proper normalization and interfaces.
  • Grasp the mathematical basis of Semantic Search (Vector Embeddings and Cosine Similarity).

Further Reading:

  • Explore the Ontology SDK documentation for implementing custom front-end applications.
  • Review Action Type security configurations to understand how granular permissions are enforced during write-backs.
  • Study the Foundry Functions documentation to learn how to write performant TypeScript for derived properties.
Ontology-Aware Applications and Logic - Ontology building - diagram 1
Ontology-Aware Applications and Logic - Ontology building - diagram 1

Semantic Search and AI Integration

Key concepts: Semantic Search · Embeddings & Vectors · Chunking Strategy · Ontology-augmented generation (RAG) · K-Nearest Neighbors (KNN) · Pix2Struct

Implementing advanced search workflows using embeddings, vector properties, and RAG.

Semantic Search and AI Integration

The modern enterprise is no longer starved for data; it is starved for context. While the Palantir Ontology serves as the "operational layer"—a high-fidelity digital twin mapping integrated data assets to real-world entities—the vast majority of organizational knowledge remains trapped in unstructured formats: PDFs, technical manuals, shift logs, and images.

Semantic Search and AI Integration represent the connective tissue that allows the Ontology to transcend structured rows and columns. By leveraging high-dimensional vector representations and Large Language Models (LLMs), we can treat unstructured text as a first-class citizen within the Ontology. This lecture explores the mechanics of transforming raw digital assets into semantically searchable objects, the mathematical foundations of vector proximity, and the advanced architectures required for Ontology-Augmented Generation (OAG).

AI_SVGI_SVGThe Pipeline of Semantic Integration: From Raw Data to Ontology-Aware Intelligence.*


1. Embeddings and the Geometry of Meaning

At the heart of semantic search lies the concept of Embeddings. In a traditional keyword-based system (e.g., BM25 or TF-IDF), search is an exercise in lexical matching. If a user searches for "velocity," a keyword system may miss documents containing "speed" or "momentum."

What it is

An embedding is a mapping of discrete tokens (words, sentences, or images) into a continuous, high-dimensional vector space $\mathbb{R}^d$.

Definition: An embedding function $E: S \to \mathbb{R}^d$ transforms a string $S$ into a vector such that the geometric distance between $E(S_1)$ and $E(S_2)$ correlates with their semantic similarity.

Why it matters

By representing meaning as a position in space, we move from "matching strings" to "calculating intent." This allows the Ontology to handle synonyms, polysemy (words with multiple meanings), and cross-lingual retrieval natively.

How it works: Cosine Similarity

In high-dimensional spaces (often $d=1536$ for models like OpenAI's text-embedding-3-small), the magnitude of the vector is often less important than its direction. We use Cosine Similarity to measure the angle $\theta$ between two vectors $\mathbf{A}$ and $\mathbf{B}$:

$$\text{similarity} = \cos(\theta) = \frac{\mathbf{A} \cdot \mathbf{B}}{|\mathbf{A}| |\mathbf{B}|}$$

Metric Formula Use Case
Cosine Similarity $\frac{\sum A_i B_i}{\sqrt{\sum A_i^2} \sqrt{\sum B_i^2}}$ Standard for text; focuses on orientation.
Euclidean Distance (L2) $\sqrt{\sum (A_i - B_i)^2}$ Useful when vector magnitude (e.g., frequency) matters.
Dot Product $\sum A_i B_i$ Used in models where vectors are not normalized.

2. Chunking Strategy: The "Goldilocks" Problem

One cannot simply feed a 200-page technical manual into an embedding model. Most models have a context window (token limit), and embedding an entire document into a single vector results in "semantic dilution," where specific details are lost in the average of the whole.

What it is

Chunking is the process of segmenting a document into smaller, semantically coherent "atoms" before they are indexed as objects in the Ontology.

Chunking Methodologies

Choosing a strategy is a trade-off between context preservation and retrieval precision.

Strategy Description Pros Cons
Fixed-Size Split by $N$ characters or tokens. Simple, computationally cheap. Often cuts mid-sentence; loses context.
Recursive Character Splits by paragraphs, then sentences, then words. Preserves structural integrity. Variable chunk sizes can be harder to manage.
Semantic Chunking Uses an LLM to find "natural" breaks in topic. Highest quality; maintains meaning. Computationally expensive; slow.
Sliding Window Fixed-size chunks with $X%$ overlap. Ensures context is shared across boundaries. Redundant data; higher storage costs.

Common Pitfalls: The Context Gap

If a chunk says, "This valve must be closed during emergencies," but the previous chunk identified the valve as "Valve-A4," the second chunk is useless in isolation. Contextual Header Injection—where the document title or section heading is prepended to every chunk—is a critical best practice in Palantir Pipeline Builder.


3. K-Nearest Neighbors (KNN) in the Ontology

Once documents are chunked and embedded, they are stored as Vector Properties on Ontology Objects. To retrieve them, we utilize the K-Nearest Neighbors (KNN) algorithm.

The Mechanics of Retrieval

When a user submits a query, the system:

  1. Embeds the query using the same model used for the documents.
  2. Performs a similarity search across the vector space.
  3. Returns the top $K$ objects with the highest similarity scores.

Computational Complexity

A naive search (Brute Force) has a complexity of $O(N \cdot D)$, where $N$ is the number of objects and $D$ is the dimensionality. For an Ontology with millions of flight records or sensor logs, this is untenable. Palantir utilizes Approximate Nearest Neighbor (ANN) indexes, such as HNSW (Hierarchical Navigable Small World), which reduces search time to $O(\log N)$ by creating a multi-layered graph of vectors.

AI_DEMOVisualization: Observe how a query vector traverses a Hierarchical Navigable Small World (HNSW) graph to find the nearest cluster of semantic neighbors.*


4. Ontology-Augmented Generation (RAG)

While semantic search finds documents, Ontology-Augmented Generation (OAG)—a specialized form of Retrieval-Augmented Generation (RAG)—uses those documents to ground LLM responses in organizational truth.

The OAG Workflow

  1. Retrieval: Use KNN to find relevant chunks.
  2. Augmentation: Inject these chunks into the LLM's prompt as "Context."
  3. Generation: The LLM answers the user's question only using the provided context.

Advanced Technique: HyDE (Hypothetical Document Embeddings)

Queries are often short and lack the semantic richness of the documents they seek. HyDE solves this by:

  1. Asking an LLM to generate a fake (hypothetical) answer to the query.
  2. Embedding that fake answer.
  3. Using the fake answer's vector to search the Ontology.

This works because the fake answer is semantically closer to the real answer in vector space than the original short query was.

Hybrid Search and Reciprocal Rank Fusion (RRF)

To ensure robustness, we often combine Semantic Search with Keyword Search. Reciprocal Rank Fusion (RRF) is the mathematical method used to combine these two disparate lists into a single ranked output.

Theorem (RRF): The score for a document $d$ is given by: $$RRFscore(d \in D) = \sum_{r \in R} \frac{1}{k + r(d)}$$ Where $r(d)$ is the rank of document $d$ in list $R$, and $k$ is a constant (typically 60).


5. Multimodal Integration: Pix2Struct

Modern Ontologies must process more than just text. Diagrams, schematics, and UI screenshots contain vital operational data.

What it is

Pix2Struct is a pretrained image-to-text model designed for "visually-situated language understanding." Unlike standard OCR (Optical Character Recognition), which just reads text, Pix2Struct understands the structure of the visual information.

How it works: Screenshot Parsing

Pix2Struct is trained on a unique objective: parsing masked screenshots into simplified HTML. By learning the relationship between pixels and underlying code structures, it can:

  • Extract data from complex tables.
  • Interpret flowcharts and organizational hierarchies.
  • Translate UI components into semantic descriptions.

Implementation in Palantir

In the Ontology, a Media Set containing technical diagrams can be processed via a Pix2Struct model. The resulting "HTML-like" text is then chunked and embedded, allowing a user to search for a specific component in a diagram via a natural language query.


6. Implementation: Code and Configuration

To implement this in the Palantir ecosystem, we define Action Types and Functions that interact with the vector index. Below is a high-signal example of how one might define a search function in a TypeScript-based Ontology environment.

import { Objects, Search, Vector } from "@foundry/ontology-api";

/**
 * Performs a Semantic Search across Technical Manuals
 * @param queryText The user's natural language query
 * @param limit The number of results to return
 */
export async function semanticManualSearch(queryText: string, limit: number = 5): Promise<ManualChunk[]> {
    // 1. Convert the query string into a vector using the embedded model
    const queryVector: Vector = await Models.get("text-embedding-ada-002").embed(queryText);

    // 2. Execute KNN search on the 'ManualChunk' object type
    // 'embedding' is a property of type 'Vector' defined in the Ontology Manager
    const results = await Objects.search()
        .manualChunk()
        .filter(chunk => chunk.embedding.knn(queryVector))
        .limit(limit)
        .all();

    return results;
}

Best Practices for Implementation

  • Model Consistency: Always use the exact same model (and version) for both indexing and querying. Switching from Ada-002 to Ada-003 will render your existing vector index invalid.
  • Vector Normalization: Ensure your vectors are unit-normalized if using dot-product similarity to prevent "long" documents from unfairly dominating the results.
  • Asymmetric Embedding: For short-query-to-long-document search, consider models specifically trained for asymmetric tasks (e.g., MSMARCO-based models).

7. Summary of AI Integration in the Ontology

The integration of AI into the Ontology transforms it from a static database into a dynamic reasoning engine. By mapping unstructured data into the same semantic space as structured objects, organizations can achieve a "Digital Twin" that not only knows what happened but can explain why based on the totality of its documentation.

Component Role in the Ontology Key Technology
Object Types The "Nouns" (e.g., Equipment, Manuals). OMA (Ontology Management App)
Vector Properties The "Meaning" of the object. Embeddings (Ada, BERT)
Action Types The "Verbs" (e.g., Update Status, Search). TypeScript Functions
RAG/OAG The "Reasoning" layer. LLMs (GPT-4, Claude)
Multimodal The "Vision" layer. Pix2Struct, UDOP

AI_FLASHCARDSI_FLASHCARDS Embedding: A numerical vector representing the meaning of data.

  • Chunking: Breaking text into smaller pieces for better retrieval.
  • KNN: K-Nearest Neighbors; the algorithm for finding similar vectors.
  • HyDE: Hypothetical Document Embedding; using a fake answer to improve search.
  • Pix2Struct: A model for understanding visually-situated language in images.
  • Reciprocal Rank Fusion: A method for merging keyword and semantic search results.

AI_QUIZI_QUIZ. Why is Cosine Similarity preferred over Euclidean Distance for text embeddings?

  • (Answer: It focuses on the direction/orientation of the vector rather than the magnitude, which is more representative of semantic intent than word count.)
  1. What is the primary risk of a "Fixed-Size" chunking strategy?
    • (Answer: It may split a sentence or paragraph in the middle, losing the context necessary for the LLM to understand the chunk.)
  2. How does HyDE improve retrieval for short queries?
    • (Answer: It expands the query into a full "hypothetical" document, providing more semantic overlap with the target documents in the vector space.)
  3. In the context of Pix2Struct, what does "visually-situated language" mean?
    • (Answer: It refers to text whose meaning is dependent on its visual context, such as its position in a table or a flowchart.)

AI_STUDY_GUIDEI_STUDY_GUIDE*Lecture Summary: Semantic Search & AI Integration**

  • Core Objective: To enable natural language interaction with the Palantir Ontology by converting unstructured data into structured vector representations.
  • Key Mathematical Concept: High-dimensional vector spaces and the use of Cosine Similarity to measure semantic proximity.
  • Engineering Requirement: Effective chunking strategies (recursive or semantic) are required to maintain context and stay within model token limits.
  • Advanced Retrieval: Move beyond simple KNN to Hybrid Search (Semantic + Keyword) and HyDE for improved accuracy.
  • Operationalization: Use the Ontology Management Application (OMA) to define vector properties and Action Types to trigger AI-driven workflows.
  • Future-Proofing: Incorporate multimodal models like Pix2Struct to bring diagrams and visual data into the searchable digital twin.
Semantic Search and AI Integration - Ontology building - diagram 1
Semantic Search and AI Integration - Ontology building - diagram 1

Source Materials

Study Ontology building with AI — Free on Lykke

Sign up for free to generate personalized flashcards, quizzes, and study guides from this course. Chat with an AI tutor that knows the material.

Get Started Free

View this course wiki on Lykke · Browse all public course wikis

Ontology building | Lykke Course Wiki