Verification checklist
- ✓Split model input and output tokens from per-call web search. A token total does not explain the search line.
- ✓Search code and Bailian apps for
enable_search,web_search, andforced_search. Confirm the switch is still on. - ✓On DashScope, read
search_infoandusage.plugins. On the OpenAI-compatible API, compare input tokens with the question alone. - ✓Check the region. Beijing turbo rates do not apply to agent search in Singapore.
- ✓Do not treat the MCP plaza’s first 2,000 free calls as free calls for built-in web search.
1. The search fee is not leftover tokens
A model-token total that matches the price sheet is not the Bailian bill. The web search guide splits built-in search into two charges: a strategy fee per call, and the extra input tokens created when retrieved pages are concatenated into the prompt, priced at that model’s standard input rate. Both ride the same HTTP request. They are not one line.
The billing guide puts this under “why is there a balance due if I barely used it”: add-ons such as web search bill per call, after the fact, separate from model inference. If enable_search is still in code or in an old app, calls keep generating search fees while nobody opens the console. Closing the chat window is not the switch.
A rollup of “what this model cost this month” smears the per-call fee into the token price, and someone swaps to a cheaper model. The lever is the search switch and the region, not the weights. The same model name only means the token column used one price sheet. It does not mean search stayed on that sheet.
2. The unit price is the strategy and the region
The guide says built-in web search itself has no free calls. From 00:00 on 27 February 2026, in China (Beijing), turbo is 3 yuan per 1,000 calls and max is 4 yuan per 1,000. That pair applies only to Beijing. Agent search is 4 yuan per 1,000 in Beijing, US (Virginia), Hong Kong, Japan (Tokyo), and Germany (Frankfurt). Singapore is listed at 73.392381 yuan per 1,000. The same code, pointed at Singapore, is not “a bit more expensive.” It is another order of magnitude. agent_max uses the same regional search price, page fetching is temporarily free, and that strategy is limited to thinking mode on qwen3-max and qwen3-max-2026-01-23.
The Responses API is the quiet path. The guide says search_options.search_strategy is silently ignored and does not turn search on. Whether a call searches depends on a mounted web_search tool and on the model’s own decision. On the Responses API, the web search tool bills like the agent strategy. A client that sends turbo can still be invoiced as agent. The dropped parameter does not return 400, so a review that only watches errors never sees it.
The guide suggests turbo for ordinary questions, and max or agent when you need several sources. Treating agent as “more accurate, only slightly dearer” misses the region. Singapore agent is not Beijing turbo plus a yuan. Prices follow the official page and will change. Reconciliation has to record the region of that day, not only the model name.
3. The switch can be on and this call still did not search
enable_search: true does not mean every call searches and every call pays the strategy fee. The model may decide the question needs no live information and answer from its own weights. To search every time, set forced_search: true in search_options. Forced search turns an occasional per-call fee into one fee per request, and it also pastes pages into the input every time. An eval set that forces search will move both the token bill and the per-call bill.
The rate limit is a second silent path. Web search is capped at 15 RPS on the Alibaba Cloud account, across every API key and every model, not per model. Above the cap the API does not error, and the search path does not run. An alert that only watches 4xx and 5xx stays quiet. Users see “search is on, but the answer looks offline.” That is not “search was free.” If search did not run, the response has no per-call evidence. If it ran, the strategy meter applies.
DashScope can be reconciled. A search that ran returns search_info, and usage includes plugins. If it did not run, both are absent. The OpenAI-compatible API cannot yet say from the response whether search ran. Compare input tokens with the question alone. The documented example is “Hangzhou weather”: about 10 input tokens offline, about 1,953 after search. A log that only stores enable_search records intent, not a paid search.
One failure looks like a miss. If search results trip content safety, the API returns DataInspectionFailed, HTTP 400, and a refusal. The guide does not say that 400 refunds the strategy fee. A refusal is not proof the call was not metered. The check is to send the same question with search off: if it then answers, the block came from the search result. On the compatible API there is no plugins field, so the inflated input-token count is the remaining evidence.
4. Which line free quota can actually cover
The free-quota guide says the grant deducts realtime model inference only. The web search guide says the feature itself has no free calls. Together: the per-call strategy fee is outside the new-user grant. The extra input tokens from pasted pages are still standard model input, so that column can still be covered by quota, resource packs, or a savings plan. “We still have free tokens” explains the token column. It does not zero the per-call column.
MCP plaza web search is a second ledger. The same web search page says every account gets 2,000 free MCP calls, then 29 yuan per 1,000, and that this does not offset built-in search. Kimi models on Bailian can use enable_search or an agent tool such as bailian_web_search. Both are called web search. They are not one invoice. Spending the MCP free calls as if they were free turbo calls leaves a gap every month.
An arrears banner is also not proof the model tokens ran over. The billing guide lists two causes: enable_search is still being called, or another pay-as-you-go product on the same Alibaba Cloud account, such as ECS or OSS, has pushed available credit negative. Credit is account-wide. Read the product on the bill before turning search off or paying a different product. Pay-as-you-go is reserve-then-settle-monthly: a posted line is not cash already taken. Model inference bills usually appear 2 to 10 minutes after the call, so a stopped client can still receive a late line, and the arrears figure can keep moving after you think you stopped.
5. Split the ledgers. A new card does not turn search off
Keep four columns: enable_search and search_strategy on the request, the region actually called, search_info or plugins on the response (or the input-token gap on the compatible API), and the per-call line that is not model tokens. When they disagree, trust the response and the bill. turbo in, agent out, on the Responses API, is a dropped parameter, not a missing charge. HTTP 200 with no plugins means search did not run, not that search was free.
Search fees leave the Alibaba Cloud balance. They do not leave a card BIN. A new virtual card does not delete enable_search, and it does not change the Singapore versus Beijing per-call gap. Web-seat renewals can still be isolated from cloud top-ups. A capped virtual card such as RDVCC can pay the seat so that renewal does not share a card with API credit. The card does not stop the silent 15 RPS skip, and it does not merge the strategy fee into the token line.