Yworq AI Safety & Responsible Use

Our commitment to building AI tools that are powerful, useful, and safe.

Our Commitment to Responsible AI

At Yworq, we are committed to building AI tools that are powerful, useful, and safe. Our platform is designed to support creative and professional productivity while actively preventing misuse.

We implement multiple layers of safeguards to ensure that our systems cannot be used to generate illegal, harmful, or inappropriate content.

Our safety architecture combines technical guardrails, policy enforcement, and continuous monitoring to protect users and the broader digital ecosystem.

Our terms and conditions also explain in detail here.

Prohibited Content

Yworq strictly prohibits the generation or use of the platform for content that includes:

  • Illegal content of any kind
  • Explicit sexual or pornographic material
  • Non-consensual or exploitative imagery
  • Content involving minors in sexual contexts
  • Violent or harmful imagery
  • Harassment or abusive content
  • Fraud, scams, or deceptive practices
  • Content that violates applicable laws or regulations

Any attempt to generate prohibited material is automatically blocked. Accounts that attempt to misuse the platform may be suspended or permanently banned.

AI Guardrails and Safety Controls

Our systems include multiple layers of automated safeguards designed to prevent harmful content generation.

  • Prompt Filtering: User prompts are evaluated before processing. Requests that contain unsafe or prohibited content are rejected.
  • Output Moderation: AI-generated content is monitored to ensure outputs comply with our safety standards and acceptable use policies.
  • Policy Enforcement: Content that violates platform policies is automatically blocked.
  • Continuous Improvement: Our safety systems are continuously updated to address emerging risks and improve content filtering capabilities.

Abuse Prevention

Yworq actively works to prevent misuse of its AI tools. We monitor for patterns of suspicious activity and implement controls that prevent attempts to bypass safety systems. Users who attempt to circumvent safeguards may face account suspension or termination.

Responsible AI Development

Our AI development process emphasizes responsible deployment of generative technologies. We focus on building systems that:

  • Support creative and professional work
  • Prevent harmful or illegal uses
  • Protect users and the public
  • Maintain transparency and accountability

Safety and compliance are core principles guiding the evolution of the Yworq platform.

Compliance with Payment Network Policies

Yworq enforces strict safeguards to prevent the generation of prohibited content, including explicit adult material and illegal imagery. Our platform is designed for legitimate creative, business, and productivity use cases and does not permit content that violates applicable payment network rules or legal standards.

System Flow & Architecture

The following diagram illustrates the lifecycle of a single user request flowing through our secure pipeline.

Loading diagram…

Component Deep-Dive

1. Pre-Generation Security (Input Guardrails)

Before we dedicate expensive GPU compute to generating an image, we inspect the user's intent. This component utilizes heuristic regular expressions against known "jailbreak" patterns and an ML-based text classifier (ezb/NSFW-Prompt-Detector) to evaluate the contextual safety of the prompt.

The Analogy: Think of this as the highly trained Bouncer at the front door of our application. If a user tries to sneak in a disguised request for prohibited content, the bouncer denies entry before the user ever gets to interact with the engine.

Business Value: Prevents the platform from generating CSAM, extreme violence, or targeted harassment, dramatically reducing legal liability.

2. The Engine

This is the open-source model (e.g., Stable Diffusion) that converts text to pixels. Because we fully encapsulate this model within our proprietary pipeline, the user never interacts with it directly.

The Analogy: The Printing Press. It does exactly what it is told, which is why the Bouncer (Guardrails) and the Notary (C2PA) are required to bookend it.

3. DWT-DCT Invisible Watermarking (Pixel Provenance)

Once the image is generated, we embed a platform-specific identifier (e.g., "Yworq AI") directly into the frequency domain (Discrete Wavelet/Cosine Transform) of the image's pixels.

The Analogy: Imagine weaving microscopic, invisible DNA strands directly into the fibers of a sheet of paper. Even if someone tears the paper in half (cropping) or crumples it up (compression and low-quality sharing on WhatsApp), the DNA remains intact and can be extracted in a forensic lab.

Business Value: If an image goes viral and metadata is stripped by social media platforms, we can mathematically prove whether it originated from our platform or a competitor's.

4. IPTC & EXIF Metadata (Standardized Transparency)

We inject standard Adobe XMP and IPTC (International Press Telecommunications Council) tags into the fileheader to declare http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia.

The Analogy: Writing a clear Return Address and Name Tag on the outside of an envelope. It tells search engines (like Google) and news agencies immediately what they are looking at.

Business Value: Ensures automated compliance with search engine guidelines and news publication standards, avoiding our user's content being blacklisted as deceptive.

5. C2PA Cryptographic Signing

This is the ultimate seal of authenticity. Created by the Coalition for Content Provenance and Authenticity (C2PA), this protocol hashes the generated pixels along with the "AI Generated" metadata and cryptographically signs it using our private RSA/ES256 keys. We define our platform as the authoritative SoftwareAgent.

The Analogy: Placing the generated image into a tamper-evident plastic bag, putting a Wax Seal on it, and stamping it with a ring that only we possess.

Business Value: If a user attempts to alter the image or erase the "AI-generated" metadata, the cryptographic seal mathematically breaks. Downstream platforms (like TikTok, X, or YouTube) will automatically know the image was tampered with after leaving our platform. It shifts the liability entirely off our shoulders.

Hazard Taxonomy and Policy

Yworq is designed to predict safety labels on the 13 categories shown below, based on the MLCommons taxonomy of 13 hazards.

S1 – Violent Crimes

Responses that enable, encourage, or endorse the commission of violent crimes, including: (1) unlawful violence toward people (ex: terrorism, genocide, murder, hate-motivated violent crimes, child abuse, assault, battery, kidnapping) and (2) unlawful violence toward animals (ex: animal abuse)

S2 – Non-Violent Crimes

Responses that enable, encourage, or endorse the commission of non-violent crimes, including: (1) personal crimes (ex: labor trafficking, threats, intimidation, hate-motivated non-violent crimes), (2) financial crimes (ex: fraud, scams, money laundering), (3) property crimes (ex: theft, arson, vandalism), (4) drug crimes (ex: creating or using regulated substances), (5) weapons crimes (ex: producing unlicensed firearms), and (6) cyber crimes (ex: hacking)

S3 – Sex-Related Crimes

Responses that enable, encourage, or endorse the commission of sex-related crimes, including: (1) sex trafficking, (2) sexual assault (ex: rape), (3) sexual harassment (physical, verbal, or visual), and (4) prostitution

S4 – Child Sexual Exploitation

Responses that contain, describe, enable, encourage, or endorse the sexual abuse of children

S5 – Defamation

Responses that are both verifiably false and likely to injure a living person's reputation

S6 – Specialized Advice

Responses that contain specialized financial, medical, or legal advice, or that indicate dangerous activities or objects are safe

S7 – Privacy

Responses that contain sensitive, nonpublic personal information that could undermine someone's physical, digital, or financial security

S8 – Intellectual Property

Responses that may violate the intellectual property rights of any third party

S9 – Indiscriminate Weapons

Responses that enable, encourage, or endorse the creation of indiscriminate weapons, including: (1) chemical weapons (ex: nerve gas), (2) biological weapons (ex: anthrax), (3) radiological weapons (ex: cobalt bombs), (4) nuclear weapons (ex: fission bombs), and (5) high-yield explosive weapons (ex: cluster munitions)

S10 – Hate

Responses that demean or dehumanize people on the basis of their sensitive, personal characteristics (i.e., race, color, ethnicity, national origin, disability, religious affiliation, caste, sexual orientation, sex, gender identity, and/or serious disease)

S11 – Suicide & Self-Harm

Responses that enable, encourage, or endorse acts of intentional self-harm, including: (1) suicide, (2) self-injury (ex: cutting), and (3) disordered eating

S12 – Sexual Content

Responses that contain erotica

S13 – Elections

Responses that contain factually incorrect information about electoral systems and processes, including in the time, place, or manner of voting in civic elections

Reporting Safety Concerns

If you believe the platform is being misused or have concerns about generated content, please contact us. We take all safety reports seriously and investigate potential violations promptly.

Email: digital@yworq.com