IT News
Apple chce vstoupit do serverů. Technologie by však musel koupit od Nvidie
Známe nový hlas Poldy Pankráce. Role české adventurní legendy se zhostí herec Pavel Liška
Česko má v čipech slušnou pozici. Z příchodu TSMC do Evropy může hodně těžit, říká exšéf tchajwanského gigantu
USA mají zbraně na oběžné dráze. Poprvé to někdo řekl nahlas, a tím to skončilo
Ceny GeForce RTX 5090 vyskočily na ~$8000, na výběr je v segmentu za ~$10 000
Hry zadarmo, nebo se slevou: Série Sniper Elite za pár stovek a rytířská klání For Honor zdarma
Hypotéky pro OSVČ: Banky umí posuzovat příjmy lépe než před lety, podmínky úvěrů se ale výrazně liší
Odškodnění za pracovní úraz: Kdy zaměstnavatel nemusí platit a co vám musí prokázat
Spuštění Intel 14A nebylo urychleno na první kvartál 2027, jde o starý údaj
LLMs respond differently to harmful prompts when AI watermarking is used
In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text, an approach Google created and released as open source. It uses a secret key that subtly changes the process a model uses for choosing the next word in a sentence. Whereas a top next word choice might be “cloudy,” the key might change it to “overcast.” Anyone who knows the key can determine if it was generated by the platform using it.
New research shows that SynthID-Text can change not just word selection but also the tools a model invokes and the chances it will adhere to or disregard safety guardrails it has been trained to follow. The threat can become greater in the face of an adversarial prompt, in which an attacker attempts to cause a model to carry out a harmful action, such as revealing a password or other sensitive information. Instructions that normally wouldn’t be followed will, in some cases, be performed once the watermarking is deployed. The finding underscores the need for developers to thoroughly test how their LLMs and agents behave when watermarking is in place.
Changing safety behavior“As compared to the same models without watermarking, it is definitely going to change their behavior, especially when we place it under adversarial conditions, or we make these models call tools when they’re powering an agent,” Andrea Siposova, an AI security researcher at Lasso Security, told Ars. “Watermarking is made to not be perceptible to a reader, but we know that when we are changing anything about what the model is generating, it is going to cause some tradeoffs, it’s going to show up somewhere.”
Pět důvodů, proč si nekupovat chytré hodinky. A dva důležité, proč ano
30 nejlepších filmů a seriálů o sériových vrazích. Psychopati, kteří děsí i baví
Balíčky z Číny zdraží od listopadu 2026. Češi si zřejmě připlatí 50 Kč za každou položku v objednávce
Googlebooky jsou za rohem. Google vyrazí proti Windows a macOS s upraveným Androidem
DJI Osmo 360 II je vlajková akční kamera. Má čtvercový snímač, vyměnitelné objektivy a šílené parametry
Těhotná želvuška, hojící se rána a divoké bitvy pestřenek. Nikon vyhlásil nejlepší videa z mikroskopů za rok 2026
Srovnávací test tiskáren na fotky – tentokrát bez pořadí, každá je úplně jiná. A někdy se vyplatí prostě zajít do fotolabu...
Druhý blok jaderné elektrárny Temelín se zastavil. Za vše mohla vadná karta
Co se teď nejvíc hraje na Nintendu Switch. Nejoblíbenější hry pro starou i novou konzoli
Apple by se mohl vrátit do serverů. Šušká se, že něco velkého peče společně s Nvidií
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- 9
- …
- následující ›
- poslední »



