This is a Jupyter notebook

Observability for Z.ai with Langfuse

This guide shows you how to integrate Langfuse with Z.ai. Z.ai's API endpoints are fully compatible with the OpenAI SDK, allowing you to trace and monitor your AI applications seamlessly.

What is Z.ai? Z.ai provides API access to the GLM family of large language models (such as GLM-4.6) through OpenAI-compatible endpoints.

What is Langfuse? Langfuse is an open-source LLM engineering platform that helps teams trace API calls, monitor performance, and debug issues in their AI applications.

Step 1: Install Dependencies

%pip install openai langfuse

Step 2: Set Up Environment Variables

Get your Langfuse keys from the project settings in Langfuse Cloud or set up self-hosting.

import os

# Get keys for your project from the project settings page: https://langfuse.com/cloud
os.environ.setdefault("LANGFUSE_PUBLIC_KEY", "pk-lf-...")
os.environ.setdefault("LANGFUSE_SECRET_KEY", "sk-lf-...")
os.environ.setdefault("LANGFUSE_BASE_URL", "https://cloud.langfuse.com") # 🇪🇺 EU region (API host)
# Other Langfuse data regions include 🇺🇸 US: https://us.cloud.langfuse.com, 🇯🇵 Japan: https://jp.cloud.langfuse.com and ⚕️ HIPAA: https://hipaa.cloud.langfuse.com

os.environ.setdefault("ZAI_API_BASE", "https://api.z.ai/api/paas/v4/")
os.environ.setdefault("ZAI_API_KEY", "...")

Step 3: Use Langfuse OpenAI Drop-in Replacement

Instead of importing openai directly, import it from langfuse.openai. Requests made through this client are automatically traced in Langfuse. Point the client at the Z.ai base URL and authenticate with your Z.ai API key.

from langfuse.openai import openai

client = openai.OpenAI(
  api_key=os.environ.get("ZAI_API_KEY"),
  base_url=os.environ.get("ZAI_API_BASE")
)

Step 4: Run an Example

response = client.chat.completions.create(
  model="glm-4.6",
  messages=[
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Why is open source important?"},
  ],
  name="Z-ai-Trace" # name of the trace
)
print(response.choices[0].message.content)

Step 5: View Traces in Langfuse

After running the example, open Langfuse Cloud to see the full trace including prompts, completions, tool calls, token usage, and latency.

Interoperability with the Python SDK

You can use this integration together with the Langfuse SDKs to add additional attributes to the observation.

The @observe() decorator provides a convenient way to automatically wrap your instrumented code and add additional attributes to the observation.

from langfuse import observe, propagate_attributes, get_client

langfuse = get_client()

@observe()
def my_llm_pipeline(input):
    # Add additional attributes (user_id, session_id, metadata, version, tags) to all spans created within this execution scope
    with propagate_attributes(
        user_id="user_123",
        session_id="session_abc",
        tags=["agent", "my-observation"],
        metadata={"email": "user@langfuse.com"},
        version="1.0.0"
    ):

        # YOUR APPLICATION CODE HERE
        result = call_llm(input)

        return result

# Run the function
my_llm_pipeline("Hi")

Learn more about using the Decorator in the Langfuse SDK instrumentation docs.

The Context Manager allows you to wrap your instrumented code using context managers (with with statements), which allows you to add additional attributes to the observation.

from langfuse import get_client, propagate_attributes

langfuse = get_client()

with langfuse.start_as_current_observation(
    as_type="span",
    name="my-observation",
    trace_context={"trace_id": "abcdef1234567890abcdef1234567890"},  # Must be 32 hex chars
) as observation:

    # Add additional attributes (user_id, session_id, metadata, version, tags)
    # to all observations created within this execution scope
    with propagate_attributes(
        user_id="user_123",
        session_id="session_abc",
        metadata={"experiment": "variant_a", "env": "prod"},
        version="1.0",
    ):
        # YOUR APPLICATION CODE HERE
        result = call_llm("some input")

# Flush events in short-lived applications
langfuse.flush()

Learn more about using the Context Manager in the Langfuse SDK instrumentation docs.

Troubleshooting

No observations appearing

First, enable debug mode in the Python SDK:

export LANGFUSE_DEBUG="True"

Then run your application and check the debug logs:

OTel observations appear in the logs: Your application is instrumented correctly but observations are not reaching Langfuse. To resolve this:
1. Call langfuse.flush() at the end of your application to ensure all observations are exported.
2. Verify that you are using the correct API keys and base URL.
No OTel spans in the logs: Your application is not instrumented correctly. Make sure the instrumentation runs before your application code.

Unwanted observations in Langfuse

The Langfuse SDK is based on OpenTelemetry. Other libraries in your application may emit OTel spans that are not relevant to you. These still count toward your billable units, so you should filter them out. See Unwanted spans in Langfuse for details.

Missing attributes

Some attributes may be stored in the metadata object of the observation rather than being mapped to the Langfuse data model. If a mapping or integration does not work as expected, please raise an issue on GitHub.

Next Steps

Once you have instrumented your code, you can manage, evaluate and debug your application:

Observability for Z.ai with Langfuse

Step 1: Install Dependencies

Step 2: Set Up Environment Variables

Step 3: Use Langfuse OpenAI Drop-in Replacement

Step 4: Run an Example

Step 5: View Traces in Langfuse

Interoperability with the Python SDK

Troubleshooting

Next Steps

Manage prompts in Langfuse

Add evaluation scores

Run LLM-as-a-judge Evaluators

Create datasets

Create custom dashboards

Test queries in the Playground

On this page