<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Tencent Cloud Log Service Trace Ingestion Cost on RockB</title><link>https://baeseokjae.github.io/tags/tencent-cloud-log-service-trace-ingestion-cost/</link><description>Recent content in Tencent Cloud Log Service Trace Ingestion Cost on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 05:41:05 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/tencent-cloud-log-service-trace-ingestion-cost/index.xml" rel="self" type="application/rss+xml"/><item><title>TencentCloud AgentObs SDK for DeepSeek Harness Review: DSH Agent Observability in 2026</title><link>https://baeseokjae.github.io/posts/tencentcloud-agentobs-sdk-dsh-observability/</link><pubDate>Thu, 01 Oct 2026 05:41:05 +0000</pubDate><guid>https://baeseokjae.github.io/posts/tencentcloud-agentobs-sdk-dsh-observability/</guid><description>Six weeks of evidence on the TencentCloud AgentObs DSH plugin: privacy defaults, real CLS cost math, adoption numbers, and when OTLP wins instead.</description><content:encoded><![CDATA[<p>For most teams evaluating dsh agent observability, the TencentCloud AgentObs SDK for DeepSeek Harness is worth adopting only if you already run Tencent Cloud Log Service. Six weeks after its first release, the evidence is mixed: solid engineering, thin adoption (14 stars, 519 npm downloads in 30 days), a privacy default that ships content capture ON, and a zero-collector architecture that replaces infrastructure you operate with metered consumption you must now watch. If you are not already a CLS customer, the community OpenTelemetry plugin is the safer default.</p>
<p>That is the short answer. This review is the third act in a series: we covered the <a href="/posts/opentelemetry-tracing-for-deepseek-harness">OpenTelemetry tracing route for DeepSeek Harness</a> in August 2026, and we published the <a href="/posts/tencentcloud-agentobs-dsh-genai-traces">installation walkthrough for this specific plugin</a> on 20 August 2026. This post does not repeat either. It is a source-backed judgement written after six weeks of releases, download telemetry and price sheets, and it is aimed at the decision rather than the setup.</p>
<h2 id="review-scope-what-six-weeks-of-evidence-changed-about-this-plugin">Review Scope: What Six Weeks of Evidence Changed About This Plugin</h2>
<p>When a first-party SDK appears, the honest question is not &ldquo;does it work&rdquo; but &ldquo;has it been adopted, and is the vendor still behind it.&rdquo; Three data points answer that.</p>
<table>
  <thead>
      <tr>
          <th>Signal</th>
          <th>TencentCloud AgentObs DSH</th>
          <th>@loongsuite/dsh-plugin</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>GitHub stars</td>
          <td>14</td>
          <td>26</td>
      </tr>
      <tr>
          <td>Forks</td>
          <td>3</td>
          <td>4</td>
      </tr>
      <tr>
          <td>Open issues</td>
          <td>0</td>
          <td>5</td>
      </tr>
      <tr>
          <td>Latest release</td>
          <td>v0.0.1 (18 Aug 2026)</td>
          <td>v0.1.2 (25 Aug 2026)</td>
      </tr>
      <tr>
          <td>Last push</td>
          <td>26 Aug 2026</td>
          <td>active</td>
      </tr>
      <tr>
          <td>npm downloads, 30 days</td>
          <td>519</td>
          <td>1,978</td>
      </tr>
      <tr>
          <td>npm downloads, 7 days</td>
          <td>121</td>
          <td>789</td>
      </tr>
      <tr>
          <td>License</td>
          <td>Apache-2.0</td>
          <td>Apache-2.0</td>
      </tr>
  </tbody>
</table>
<p>The OpenTelemetry competitor draws roughly <strong>3.8x the monthly downloads and 6.5x the weekly downloads</strong> of the Tencent plugin. Weekly-to-monthly ratio matters here: 789 of 1,978 (40%) downloads landed in the last week for the OTel plugin, versus 121 of 519 (23%) for Tencent&rsquo;s. The community plugin is accelerating; the first-party one is flattening.</p>
<p>None of that is a verdict on code quality. The Tencent plugin is at 0.1.0, six weeks old, with zero open issues — which in a low-adoption repository often means nobody has filed one rather than that nothing is wrong. Read the numbers as maturity evidence, not as a scoreboard.</p>
<p>There is also a version-hygiene detail worth noting: the GitHub release tag is <strong>v0.0.1</strong> while the npm package&rsquo;s latest version is <strong>0.1.0</strong>. Both were first published on 18 August 2026; the npm package was last modified 26 August 2026. If your upgrade tooling reads npm and your changelog reads GitHub releases, those two streams are telling you different stories.</p>
<h2 id="what-the-agentobs-dsh-plugin-does-short-recap">What the AgentObs DSH Plugin Does (Short Recap)</h2>
<p>The plugin observes DeepSeek Harness&rsquo;s native session, agent-loop, LLM-stream and tool-lifecycle events and converts them into Tencent Cloud&rsquo;s five-layer span model — entry, agent, step, chat, tool. It ships spans as Protobuf directly to Tencent Cloud Log Service through <code>tencentcloud-cls-sdk-js</code>, with no OTLP collector and no sidecar.</p>
<p>If you need the install steps, the profile-add command, the <code>pnpm approve-builds</code> workaround, the full configuration table or the span-nesting diagram, they are all in the <a href="/posts/tencentcloud-agentobs-dsh-genai-traces">original setup guide</a>. What follows assumes you have read that and are deciding whether to keep it.</p>
<h2 id="adoption-check-stars-npm-downloads-and-release-cadence-against-the-competition">Adoption Check: Stars, npm Downloads and Release Cadence Against the Competition</h2>
<p>Two details in the table above deserve separating from the raw counts.</p>
<p>First, <strong>0 open issues is not the same as 0 defects</strong>. The loongsuite plugin&rsquo;s five open issues are a sign of an active user base stress-testing edge cases — retries, aborts, content-capture boundaries. A repository with 519 monthly downloads and no issue tracker activity has not yet been pushed hard. Treat the Tencent plugin&rsquo;s bug surface as <em>unknown</em> rather than <em>clean</em>.</p>
<p>Second, <strong>download counts measure installs, not retention</strong>. 519 downloads in 30 days for a zero-collector plugin whose headline promise is &ldquo;skip the collector&rdquo; is modest. If that promise were the decisive advantage it sounds like, you would expect the ease-of-adoption curve to favour Tencent&rsquo;s plugin, not the community one — especially given Tencent&rsquo;s plugin supports Node.js &gt;=18.0.0 while the loongsuite plugin requires Node.js &gt;=22.19.0. The plugin that runs on more machines is being installed on fewer. That gap is the most interesting number in this review, and it points at the harness-ecosystem question rather than at Tencent&rsquo;s engineering.</p>
<p>The most likely explanation is stack gravity: DeepSeek Harness users are largely picking their observability backend first (Jaeger, Tempo, SigNoz, Langfuse) and their ingestion plugin second. A backend-native plugin only wins when the backend is already decided.</p>
<h2 id="privacy-defaults-compared-capturecontent-true-vs-off">Privacy Defaults Compared: captureContent True vs Off</h2>
<p>This is the section to read before anything else, because it is the only difference in this comparison that can cause an incident rather than an inconvenience.</p>
<table>
  <thead>
      <tr>
          <th>Behaviour</th>
          <th>TencentCloud AgentObs DSH</th>
          <th>@loongsuite/dsh-plugin</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Content capture default</td>
          <td><strong>ON</strong> (<code>captureContent: true</code>)</td>
          <td><strong>OFF</strong></td>
      </tr>
      <tr>
          <td>Content cap</td>
          <td><code>contentMaxChars</code> 128,000</td>
          <td>applies when enabled</td>
      </tr>
      <tr>
          <td>What gets attached</td>
          <td>prompts, responses, tool arguments, tool results</td>
          <td>same fields, when enabled</td>
      </tr>
      <tr>
          <td>Destination</td>
          <td>your CLS topic</td>
          <td>your OTLP backend</td>
      </tr>
  </tbody>
</table>
<p>A DeepSeek Harness coding session is not a chat session. It reads source trees, <code>.env</code> files, CI configuration, infrastructure manifests, database dumps and arbitrary tool output. With content capture on by default and a 128,000-character per-item cap, a single traced turn can carry a substantial slice of a private repository into a hosted log service — at 128K characters, that is on the order of 30,000 tokens of payload per captured item, and tool results can be several such items per step.</p>
<p>The defence is one configuration line. The problem is that defaults are what actually ship. In a plugin shipped by a cloud vendor to that vendor&rsquo;s own log service, default-on capture is a defensible product decision — the vendor&rsquo;s console demo depends on having content to display. But it means the <em>installation</em> is the compliance decision, not a later checkbox. If you roll this out to a fleet by profile, every developer inherits a capture-on posture unless the profile explicitly says otherwise.</p>
<p>Two mitigations, in order of preference:</p>
<ol>
<li>Set <code>captureContent: false</code> in the profile before rollout, and enable it per-developer only when debugging. This also cuts payload volume by roughly an order of magnitude — see the cost section for what that is worth.</li>
<li>If you need content, lower <code>contentMaxChars</code> aggressively (a few thousand characters, not 128,000), scope the CLS topic&rsquo;s retention, and treat the topic as you would a secrets-adjacent data store.</li>
</ol>
<p>The OTel competitor&rsquo;s off-by-default posture is the correct default for a coding agent, and it is the single strongest argument in this review for choosing it.</p>
<h2 id="what-it-actually-costs-cls-ingestion-math-for-a-realistic-dsh-workload">What It Actually Costs: CLS Ingestion Math for a Realistic DSH Workload</h2>
<p>&ldquo;Zero collector&rdquo; removes infrastructure, not cost. It moves the spend from a server you run to consumption you meter, and in CLS the meter has several line items that behave differently.</p>
<p>Chinese mainland pay-as-you-go list prices from <a href="https://www.tencentcloud.com/pricing/cls">Tencent Cloud&rsquo;s CLS pricing page</a>:</p>
<table>
  <thead>
      <tr>
          <th>Line item</th>
          <th>Rate</th>
          <th>Billed on</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Log write traffic</td>
          <td>USD 0.032 / GB / day</td>
          <td><strong>compressed</strong> volume</td>
      </tr>
      <tr>
          <td>Standard index traffic</td>
          <td>USD 0.062 / GB / day</td>
          <td><strong>uncompressed</strong> volume</td>
      </tr>
      <tr>
          <td>Standard log storage</td>
          <td>USD 0.0024 / GB / day</td>
          <td><strong>compressed</strong> volume</td>
      </tr>
      <tr>
          <td>Standard index storage</td>
          <td>USD 0.0024 / GB / day</td>
          <td><strong>uncompressed</strong> volume</td>
      </tr>
      <tr>
          <td>Data processing</td>
          <td>USD 0.026 / GB / day</td>
          <td>compressed volume</td>
      </tr>
      <tr>
          <td>Partition</td>
          <td>USD 0.007 / partition / day</td>
          <td>fixed</td>
      </tr>
      <tr>
          <td>Service requests</td>
          <td>USD 0.026 / million / day</td>
          <td>count</td>
      </tr>
  </tbody>
</table>
<p>The asymmetry is the whole story: <strong>index traffic costs roughly 2x write traffic and is charged on uncompressed bytes</strong>, while write traffic and storage are charged on compressed bytes (typical log compression runs 1:4 to 1:10). Tencent&rsquo;s own worked billing examples show index traffic and index storage as the two largest line items. Your bill follows <em>how many span fields you index</em>, not how many traces you emit.</p>
<p>Working through a realistic DSH fleet, assuming 58 spans per agent task, 1:4 compression, one partition, and 35% of raw volume carrying indexed fields:</p>
<table>
  <thead>
      <tr>
          <th>Daily agent tasks</th>
          <th>Raw volume</th>
          <th>Variable cost / month</th>
          <th>Partition / month</th>
          <th>Total / month</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>10</td>
          <td>17 MB/day</td>
          <td>$0.02</td>
          <td>$0.21</td>
          <td>$0.23</td>
      </tr>
      <tr>
          <td>25</td>
          <td>43 MB/day</td>
          <td>$0.05</td>
          <td>$0.21</td>
          <td>$0.26</td>
      </tr>
      <tr>
          <td>100</td>
          <td>170 MB/day</td>
          <td>$0.20</td>
          <td>$0.21</td>
          <td>$0.41</td>
      </tr>
      <tr>
          <td>1,000</td>
          <td>1.7 GB/day</td>
          <td>$1.96</td>
          <td>$0.21</td>
          <td>$2.17</td>
      </tr>
      <tr>
          <td>5,000</td>
          <td>8.5 GB/day</td>
          <td>$9.82</td>
          <td>$0.21</td>
          <td>$10.03</td>
      </tr>
      <tr>
          <td>20,000</td>
          <td>34 GB/day</td>
          <td>$39.29</td>
          <td>$0.21</td>
          <td>$39.50</td>
      </tr>
  </tbody>
</table>
<p>Two conclusions fall out of this and both are counter-intuitive.</p>
<p><strong>At small scale, the fixed partition charge dominates and the whole thing is nearly free.</strong> A team running 10 to 100 agent tasks a day pays well under a dollar a month. Anyone running a self-hosted Jaeger or Tempo instance to avoid that is spending far more on the VM than the metered alternative costs.</p>
<p><strong>At large scale, content capture — not trace count — sets the bill.</strong> Recomputing the 20,000-task-per-day case with span payloads of 30 KB instead of 8 KB (which is what capture-on looks like once prompts and tool results are attached):</p>
<ul>
<li>index 0% of volume: <strong>$16.84/month</strong></li>
<li>index 10%: <strong>$23.26/month</strong></li>
<li>index 35%: <strong>$39.29/month</strong></li>
<li>index 60%: <strong>$55.32/month</strong></li>
<li>index 100%: <strong>$80.96/month</strong></li>
</ul>
<p>That is a 4.8x spread on the same traces, driven entirely by index configuration — and the index ratio is the lever most teams never touch. Combined with the capture-on/off comparison at 1,000 tasks per day (about $1.96/month with capture on versus $0.22/month with it off), disabling content capture is simultaneously the strongest privacy control and roughly a 9x cost reduction.</p>
<p>Where self-hosting wins is a narrow band. A $12/month VM running Jaeger or Tempo has near-zero marginal ingestion cost, which beats CLS somewhere above 5,000 to 20,000 agent tasks per day depending on payload size. Below that, the arithmetic favours CLS unless you already own the infrastructure for other reasons. Above it, the arithmetic still favours CLS if you have turned indexing down — at 0% indexing the same 20,000-task workload is $16.84/month against $12/month plus your operational time.</p>
<p>Resource packs are denominated in U, where 1U = CNY 1, and can offset any billable item. A &ldquo;billed on raw log volume&rdquo; mode exists but is whitelist-only, so do not plan around it.</p>
<h2 id="transport-showdown-protobuf-to-cls-versus-standard-otlp">Transport Showdown: Protobuf-to-CLS Versus Standard OTLP</h2>
<p>The transport choice is where portability is decided, and it is decided against you.</p>
<table>
  <thead>
      <tr>
          <th>Dimension</th>
          <th>AgentObs DSH</th>
          <th>loongsuite dsh-plugin</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Wire format</td>
          <td>Protobuf, CLS-specific</td>
          <td>OTLP/HTTP protobuf</td>
      </tr>
      <tr>
          <td>Destination</td>
          <td>Tencent Cloud Log Service topic</td>
          <td>any OTLP backend</td>
      </tr>
      <tr>
          <td>Backends reachable</td>
          <td>CLS</td>
          <td>Jaeger, Tempo, SigNoz, Langfuse, and others</td>
      </tr>
      <tr>
          <td>Auth</td>
          <td>SecretId/SecretKey (strong) or numeric UIN (weak)</td>
          <td>backend-defined</td>
      </tr>
      <tr>
          <td>Metrics</td>
          <td>traces documented</td>
          <td><code>gen_ai.client.operation.duration</code>, <code>gen_ai.client.token.usage</code></td>
      </tr>
      <tr>
          <td>Vendor dependency</td>
          <td><code>tencentcloud-cls-sdk-js</code></td>
          <td>OpenTelemetry SDK stack</td>
      </tr>
      <tr>
          <td>Runtime dependencies</td>
          <td>2</td>
          <td>OTel SDK stack</td>
      </tr>
  </tbody>
</table>
<p>Two structural facts matter more than the table.</p>
<p>First, <strong>the Tencent plugin&rsquo;s documented surface is traces</strong>, while the OTel plugin also exports the standard GenAI metrics (<code>gen_ai.client.operation.duration</code> and <code>gen_ai.client.token.usage</code>). Metrics are what you alert on cheaply and continuously; traces are what you open when an alert fires. A traces-only integration puts more weight on your log backend&rsquo;s query and dashboard features.</p>
<p>Second, <strong>the auth modes are a real operational split</strong>. Strong auth via <code>CLS_SECRET_ID</code> and <code>CLS_SECRET_KEY</code> is the right choice, but it is a long-lived credential that the plugin holds and that will one day expire or get rotated out from under a running fleet. Weak auth with a numeric UIN and no key removes the secret but weakens the trust boundary. Neither mode gives you the short-lived, workload-identity-style credentials you would want for a fleet. Plan the rotation before you plan the rollout.</p>
<p>The portability question reduces to: <strong>can I replay the same trace into another backend?</strong> The answer for this plugin is no, not without rewriting the export path. The CLS-bound field mapping and the Protobuf-to-CLS transport are what leave with your data. That is a fair trade for a CLS-native team and a poor one for everyone else.</p>
<h2 id="the-one-backend-rule-why-you-cannot-run-this-and-the-opentelemetry-plugin-together">The One-Backend Rule: Why You Cannot Run This and the OpenTelemetry Plugin Together</h2>
<p>This is the constraint that changes the shape of the decision, and it is frequently missed.</p>
<p>DeepSeek Harness ships a public telemetry seam, <code>@deepseek-ai/dsh-session-telemetry</code> (v0.0.1-rc.1, published 10 August 2026, BSD-3-Clause), described as &ldquo;session-event capture, projection, redaction, and handoff to a reporting backend.&rdquo; The seam accepts <strong>exactly one backend per context</strong>. Loading a duplicate — the official OTLP-logs exporter plus a tracing plugin, or two tracing plugins at once — throws an error at load time.</p>
<p>The practical consequences:</p>
<ul>
<li>DSH tracing plugins are <strong>mutually exclusive, not stackable</strong>. You cannot run the Tencent plugin and the OpenTelemetry plugin side by side to compare them.</li>
<li>The harness also ships its own OTLP-logs exporter implementing the same seam, so choosing a tracing plugin means <strong>replacing</strong> the official exporter, not adding to it.</li>
<li>The choice is made <strong>at install time, per profile</strong>. It is not a runtime toggle.</li>
</ul>
<p>Because of this, a side-by-side evaluation requires two profiles, not two plugins. And it means the cost of reversing the decision is a profile edit plus a lost ingestion history: trace data already written to CLS stays in CLS, and nothing backfills into a Jaeger instance you switch to later.</p>
<p>Practical advice: if you are unsure, evaluate on a second profile first, with a bounded time box. Since you cannot run both, the migration cost is real and asymmetric — switching <em>to</em> CLS is cheap, switching <em>away</em> from it means abandoning the collected history.</p>
<h2 id="runtime-behaviour-batching-queue-drops-retries-and-silent-failures">Runtime Behaviour: Batching, Queue Drops, Retries and Silent Failures</h2>
<p>The plugin&rsquo;s buffering behaviour is documented and, on balance, sensible. It is also where you should focus your alerting.</p>
<table>
  <thead>
      <tr>
          <th>Parameter</th>
          <th>Default</th>
          <th>Consequence</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><code>batchMaxSize</code></td>
          <td>32 spans</td>
          <td>batch granularity</td>
      </tr>
      <tr>
          <td><code>maxBatchBytes</code></td>
          <td>10 MB</td>
          <td>must stay under the CLS SDK&rsquo;s 19 MB hard limit</td>
      </tr>
      <tr>
          <td><code>maxQueueSize</code></td>
          <td>2,048 spans</td>
          <td>oldest span dropped when full</td>
      </tr>
      <tr>
          <td><code>flushIntervalMs</code></td>
          <td>5,000 ms</td>
          <td>up to 5s of trace delay</td>
      </tr>
      <tr>
          <td><code>retryTimes</code></td>
          <td>3</td>
          <td>failed batch retried three times</td>
      </tr>
      <tr>
          <td><code>contentMaxChars</code></td>
          <td>128,000</td>
          <td>per captured item</td>
      </tr>
  </tbody>
</table>
<p>Three behaviours deserve reframing rather than fixing.</p>
<p><strong>Queue overflow silently drops the oldest span.</strong> At 2,048 buffered spans the plugin discards from the head of the queue to keep the harness responsive. This is a backpressure design, not a bug — and crucially, <strong>it is a signal</strong>. Steady-state dropping means your ingestion path cannot keep up with your agent workload, which is exactly the condition you want an alert on. Monitor for it rather than treating it as noise. The subtlety is that dropping the oldest span preserves recent context at the cost of the start of a long-running task — the part you often need to explain why a long task went wrong.</p>
<p><strong><code>maxBatchBytes</code> must stay below 19 MB.</strong> Exceed the CLS SDK&rsquo;s hard limit and the entire batch is rejected locally, not by the server. With capture on and a large tool result, a single oversized batch can take its sibling spans down with it. Lower the batch byte cap before you lower the content cap.</p>
<p><strong>An expired CLS key fails quietly from the harness&rsquo;s point of view.</strong> A 401 from CLS does not fail the agent run — the harness exits 0, the developer sees a normal session, and traces simply stop appearing. From the harness&rsquo;s perspective this is correct: observability must never break the workload. From an operations perspective it means <strong>&ldquo;no traces&rdquo; and &ldquo;no problems&rdquo; look identical</strong>. You need a synthetic check that asserts a known trace lands in the topic on a schedule, independent of the plugin&rsquo;s own success reporting.</p>
<p>The other two documented failure modes are installation-stage and cheap to avoid: on pnpm v9+, protobufjs build scripts are blocked by default, producing <code>ERR_PNPM_IGNORED_BUILDS</code> until you set <code>enable-scripts=true</code> and reinstall; and the plugin does not attach until the harness is restarted after installation.</p>
<h2 id="compatibility-windows-and-the-020-cliff">Compatibility Windows and the 0.2.0 Cliff</h2>
<p>Both leading plugins pin the same harness window — <code>&gt;=0.1.0-rc.6</code> and <code>&lt;0.2.0</code> — and both warn that outside that range lifecycle hooks may simply not fire. The failure mode is the nasty kind: nothing crashes, traces just stop.</p>
<table>
  <thead>
      <tr>
          <th>Constraint</th>
          <th>AgentObs DSH</th>
          <th>loongsuite dsh-plugin</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>DSH version window</td>
          <td>&gt;=0.1.0-rc.6, &lt;0.2.0</td>
          <td>&gt;=0.1.0-rc.6, &lt;0.2.0</td>
      </tr>
      <tr>
          <td>Node.js floor</td>
          <td><strong>&gt;=18.0.0</strong></td>
          <td>&gt;=22.19.0</td>
      </tr>
      <tr>
          <td>Verified on</td>
          <td>not stated publicly</td>
          <td>0.1.0-rc.6 headless and Web profiles</td>
      </tr>
  </tbody>
</table>
<p>The <strong>Node.js floor is the most underrated difference</strong>. Tencent&rsquo;s plugin supports Node &gt;=18, three major versions lower than the OTel plugin&rsquo;s &gt;=22.19.0. For a fleet still running Node 20 on older build images, this is the difference between a same-day rollout and a toolchain upgrade project. It is a genuine adoption advantage and it is the strongest technical argument for the Tencent plugin.</p>
<p>The 0.2.0 cliff is shared and asymmetric in effect. Both plugins break at the same boundary, but the recovery paths differ: the OTel plugin&rsquo;s users can pin the harness and keep exporting to a backend they control, while CLS users are dependent on Tencent shipping a 0.2.0-compatible update. Given the release cadence observed here — one tag in six weeks, no push since 26 August — that dependency carries schedule risk. Pin your harness version and treat the DSH upgrade as a change that requires an observability regression test.</p>
<h2 id="feature-by-feature-agentobs-vs-loongsuite-vs-the-langfuse-plugin">Feature-by-Feature: AgentObs vs loongsuite vs the Langfuse Plugin</h2>
<p>A third option is worth including, because it represents a different philosophy rather than a different vendor.</p>
<table>
  <thead>
      <tr>
          <th>Capability</th>
          <th>TencentCloud AgentObs DSH</th>
          <th>@loongsuite/dsh-plugin</th>
          <th>dsh-plugin-langfuse</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Version</td>
          <td>0.1.0 (npm)</td>
          <td>v0.1.2</td>
          <td>v0.7.0</td>
      </tr>
      <tr>
          <td>Philosophy</td>
          <td>backend-native</td>
          <td>backend-neutral</td>
          <td>LLM-product-team</td>
      </tr>
      <tr>
          <td>Transport</td>
          <td>Protobuf to CLS</td>
          <td>OTLP/HTTP</td>
          <td>Langfuse SDK via the telemetry seam</td>
      </tr>
      <tr>
          <td>Span shape</td>
          <td>ENTRY → AGENT → STEP → CHAT/TOOL</td>
          <td>ENTRY → AGENT → STEP → LLM/TOOL</td>
          <td>step → generation, tool call → tool span</td>
      </tr>
      <tr>
          <td>Per-retry LLM span</td>
          <td>not documented</td>
          <td>yes</td>
          <td>not documented</td>
      </tr>
      <tr>
          <td>Error/abort closes spans</td>
          <td>not documented</td>
          <td>yes</td>
          <td>not documented</td>
      </tr>
      <tr>
          <td>Metrics exported</td>
          <td>not documented</td>
          <td><code>gen_ai.*</code> standard metrics</td>
          <td>Langfuse Scores for feedback</td>
      </tr>
      <tr>
          <td>Content capture default</td>
          <td>ON</td>
          <td>OFF</td>
          <td>backend-configured</td>
      </tr>
      <tr>
          <td>Session grouping</td>
          <td>via CLS topic</td>
          <td>via OTLP resource attributes</td>
          <td>explicit, per session</td>
      </tr>
      <tr>
          <td>Subagent/fork lineage</td>
          <td>not documented</td>
          <td>not documented</td>
          <td>preserved</td>
      </tr>
      <tr>
          <td>Node.js floor</td>
          <td>&gt;=18.0.0</td>
          <td>&gt;=22.19.0</td>
          <td>not stated</td>
      </tr>
      <tr>
          <td>Backends reachable</td>
          <td>CLS only</td>
          <td>Jaeger, Tempo, SigNoz, Langfuse</td>
          <td>Langfuse</td>
      </tr>
  </tbody>
</table>
<p>Three capability gaps stand out against the OTel plugin, all of which matter for debugging agent behaviour rather than for basic tracing.</p>
<p><strong>Per-retry LLM spans.</strong> If a model call fails and retries within the same step, the OTel plugin emits a distinct LLM span per real attempt, so you can see that three calls happened and one succeeded. Whether the Tencent plugin does the same is not documented — which means you should not assume you will be able to distinguish &ldquo;one slow model call&rdquo; from &ldquo;three retried calls&rdquo; in CLS without testing it yourself.</p>
<p><strong>Error and abort handling.</strong> The OTel plugin closes spans with error status rather than leaving them open. Left-open spans are the classic reason a trace view shows an agent turn as still running hours after it died.</p>
<p><strong>Correlation of tool calls to results</strong> by DSH call ID is explicitly documented for the OTel plugin. Tool-call correlation is the single most useful feature of coding-agent tracing, because it tells you <em>which</em> tool call consumed the time.</p>
<p>The Langfuse plugin occupies its own niche: it groups traces by session, records canonical feedback as Langfuse Scores, and preserves fork and subagent lineage — capabilities neither of the other two advertise. For teams whose observability questions are about output quality and iteration rather than infrastructure, it is the closest fit to the question being asked.</p>
<p>The honest summary: the Tencent plugin competes on deployment simplicity and Node.js reach; the OTel plugin competes on trace semantics and portability; the Langfuse plugin competes on session-level product analytics.</p>
<h2 id="who-should-use-the-tencentcloud-agentobs-sdk-in-october-2026">Who Should Use the TencentCloud AgentObs SDK in October 2026</h2>
<p>There is no single winner, so here is the decision matrix instead of a verdict.</p>
<p><strong>Use it if:</strong></p>
<ul>
<li>You are already a Tencent Cloud customer with existing CLS topics, retention policies and alerting. Traces landing in the same console as the rest of your logs makes cost correlation and retention management free instead of another silo.</li>
<li>Your fleet runs Node.js 18 or 20. The Node.js floor is the clearest technical win in this comparison.</li>
<li>Your agent volume is modest — under a few thousand tasks a day — where the bill is a few dollars a month and a self-hosted stack is disproportionate overhead.</li>
<li>Your organisation has already decided that DSH traces belong in CLS, which is a stack decision above the plugin.</li>
</ul>
<p><strong>Skip it if:</strong></p>
<ul>
<li>You are not on Tencent Cloud. You would be buying lock-in to save one deployment step that SigNoz or Langfuse already match.</li>
<li>You need content capture off by default across a fleet without managing a profile override. The OTel plugin&rsquo;s default is correct for a coding agent; this one&rsquo;s is not.</li>
<li>You depend on documented per-retry spans, error-closed spans and tool-call correlation. The Tencent plugin does not publish those semantics.</li>
<li>Your harness will move past 0.2.0 on a schedule you control. Both plugins break there, but only one leaves you free to pin and keep tracing to your own backend.</li>
<li>You need vendor-neutral metrics, or the ability to replay one trace into two backends.</li>
</ul>
<p><strong>The migration cost to weigh before deciding:</strong> because of the one-backend rule, you cannot hedge. Switching to CLS is cheap; switching away abandons the history you have already written. If there is any chance your observability strategy moves toward a neutral backend, start there and stay there.</p>
<h2 id="faq">FAQ</h2>
<h3 id="is-the-tencentcloud-agentobs-sdk-worth-using-for-dsh-agent-observability">Is the TencentCloud AgentObs SDK worth using for dsh agent observability?</h3>
<p>Yes, if you already run Tencent Cloud Log Service and want traces with no collector to operate. No, if you are not a CLS customer, need content capture off by default, or want vendor-neutral trace semantics. It is a strong fit inside the Tencent Cloud stack and a lock-in purchase outside it.</p>
<h3 id="does-the-tencentcloud-agentobs-sdk-capture-prompts-and-tool-output-by-default">Does the TencentCloud AgentObs SDK capture prompts and tool output by default?</h3>
<p>Yes. It ships with <code>captureContent: true</code> and a <code>contentMaxChars</code> cap of 128,000, which attaches prompts, responses, tool arguments and tool results to spans by default. The community OpenTelemetry plugin ships content capture off by default. For a coding agent that reads source trees and <code>.env</code> files, disable it in the profile before rolling out.</p>
<h3 id="how-much-does-cls-agent-observability-cost-for-a-dsh-fleet">How much does CLS Agent Observability cost for a DSH fleet?</h3>
<p>At small scale, very little: roughly $0.20 to $0.45 per month for 10 to 100 agent tasks a day, where the fixed partition charge dominates. At 20,000 tasks a day, the variable bill runs about $39 per month at a 35% index ratio, or about $17 per month with indexing disabled. Index traffic is billed on uncompressed volume at roughly 2x the write-traffic rate, so index configuration — not trace count — drives the bill.</p>
<h3 id="can-i-run-the-tencentcloud-agentobs-sdk-alongside-an-opentelemetry-tracing-plugin">Can I run the TencentCloud AgentObs SDK alongside an OpenTelemetry tracing plugin?</h3>
<p>No. The DeepSeek Harness telemetry seam <code>@deepseek-ai/dsh-session-telemetry</code> accepts exactly one backend per context and throws at load time if you register a second. The Tencent plugin, the OTel plugins and the harness&rsquo;s own OTLP-logs exporter are mutually exclusive; evaluating two of them requires two separate profiles.</p>
<h3 id="what-happens-if-my-cls-credentials-expire-while-the-plugin-is-running">What happens if my CLS credentials expire while the plugin is running?</h3>
<p>The harness does not fail. A 401 from CLS leaves the agent run successful and the process exits 0, while traces silently stop appearing. Because observability failures are deliberately non-fatal to the workload, &ldquo;no traces&rdquo; and &ldquo;no problems&rdquo; look identical — add an independent synthetic check that asserts a known trace lands in your topic on a schedule.</p>
]]></content:encoded></item></channel></rss>