Yworq AI Safety & Responsible Use
Our commitment to building AI tools that are powerful, useful, and safe.
Our Commitment to Responsible AI
At Yworq, we are committed to building AI tools that are powerful, useful, and safe. Our platform is designed to support creative and professional productivity while actively preventing misuse.
We implement multiple layers of safeguards to ensure that our systems cannot be used to generate illegal, harmful, or inappropriate content.
Our safety architecture combines technical guardrails, policy enforcement, and continuous monitoring to protect users and the broader digital ecosystem.
Our terms and conditions also explain in detail here.
Prohibited Content
Yworq strictly prohibits the generation or use of the platform for content that includes:
- Illegal content of any kind
- Explicit sexual or pornographic material
- Non-consensual or exploitative imagery
- Content involving minors in sexual contexts
- Violent or harmful imagery
- Harassment or abusive content
- Fraud, scams, or deceptive practices
- Content that violates applicable laws or regulations
Any attempt to generate prohibited material is automatically blocked. Accounts that attempt to misuse the platform may be suspended or permanently banned.
AI Guardrails and Safety Controls
Our systems include multiple layers of automated safeguards designed to prevent harmful content generation.
- Prompt Filtering: User prompts are evaluated before processing. Requests that contain unsafe or prohibited content are rejected.
- Output Moderation: AI-generated content is monitored to ensure outputs comply with our safety standards and acceptable use policies.
- Policy Enforcement: Content that violates platform policies is automatically blocked.
- Continuous Improvement: Our safety systems are continuously updated to address emerging risks and improve content filtering capabilities.
Abuse Prevention
Yworq actively works to prevent misuse of its AI tools. We monitor for patterns of suspicious activity and implement controls that prevent attempts to bypass safety systems. Users who attempt to circumvent safeguards may face account suspension or termination.
Responsible AI Development
Our AI development process emphasizes responsible deployment of generative technologies. We focus on building systems that:
- Support creative and professional work
- Prevent harmful or illegal uses
- Protect users and the public
- Maintain transparency and accountability
Safety and compliance are core principles guiding the evolution of the Yworq platform.
Compliance with Payment Network Policies
Yworq enforces strict safeguards to prevent the generation of prohibited content, including explicit adult material and illegal imagery. Our platform is designed for legitimate creative, business, and productivity use cases and does not permit content that violates applicable payment network rules or legal standards.
System Flow & Architecture
The following diagram illustrates the lifecycle of a single user request flowing through our secure pipeline.
Component Deep-Dive
1. Pre-Generation Security (Input Guardrails)
Before we dedicate expensive GPU compute to generating an image, we inspect the user's intent. This component utilizes heuristic regular expressions against known "jailbreak" patterns and an ML-based text classifier (ezb/NSFW-Prompt-Detector) to evaluate the contextual safety of the prompt.
The Analogy: Think of this as the highly trained Bouncer at the front door of our application. If a user tries to sneak in a disguised request for prohibited content, the bouncer denies entry before the user ever gets to interact with the engine.
Business Value: Prevents the platform from generating CSAM, extreme violence, or targeted harassment, dramatically reducing legal liability.
2. The Engine
This is the open-source model (e.g., Stable Diffusion) that converts text to pixels. Because we fully encapsulate this model within our proprietary pipeline, the user never interacts with it directly.
The Analogy: The Printing Press. It does exactly what it is told, which is why the Bouncer (Guardrails) and the Notary (C2PA) are required to bookend it.
3. DWT-DCT Invisible Watermarking (Pixel Provenance)
Once the image is generated, we embed a platform-specific identifier (e.g., "Yworq AI") directly into the frequency domain (Discrete Wavelet/Cosine Transform) of the image's pixels.
The Analogy: Imagine weaving microscopic, invisible DNA strands directly into the fibers of a sheet of paper. Even if someone tears the paper in half (cropping) or crumples it up (compression and low-quality sharing on WhatsApp), the DNA remains intact and can be extracted in a forensic lab.
Business Value: If an image goes viral and metadata is stripped by social media platforms, we can mathematically prove whether it originated from our platform or a competitor's.
4. IPTC & EXIF Metadata (Standardized Transparency)
We inject standard Adobe XMP and IPTC (International Press Telecommunications Council) tags into the fileheader to declare http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia.
The Analogy: Writing a clear Return Address and Name Tag on the outside of an envelope. It tells search engines (like Google) and news agencies immediately what they are looking at.
Business Value: Ensures automated compliance with search engine guidelines and news publication standards, avoiding our user's content being blacklisted as deceptive.
5. C2PA Cryptographic Signing
This is the ultimate seal of authenticity. Created by the Coalition for Content Provenance and Authenticity (C2PA), this protocol hashes the generated pixels along with the "AI Generated" metadata and cryptographically signs it using our private RSA/ES256 keys. We define our platform as the authoritative SoftwareAgent.
The Analogy: Placing the generated image into a tamper-evident plastic bag, putting a Wax Seal on it, and stamping it with a ring that only we possess.
Business Value: If a user attempts to alter the image or erase the "AI-generated" metadata, the cryptographic seal mathematically breaks. Downstream platforms (like TikTok, X, or YouTube) will automatically know the image was tampered with after leaving our platform. It shifts the liability entirely off our shoulders.
Hazard Taxonomy and Policy
Yworq is designed to predict safety labels on the 13 categories shown below, based on the MLCommons taxonomy of 13 hazards.
S1 – Violent Crimes
Responses that enable, encourage, or endorse the commission of violent crimes, including: (1) unlawful violence toward people (ex: terrorism, genocide, murder, hate-motivated violent crimes, child abuse, assault, battery, kidnapping) and (2) unlawful violence toward animals (ex: animal abuse)
S2 – Non-Violent Crimes
Responses that enable, encourage, or endorse the commission of non-violent crimes, including: (1) personal crimes (ex: labor trafficking, threats, intimidation, hate-motivated non-violent crimes), (2) financial crimes (ex: fraud, scams, money laundering), (3) property crimes (ex: theft, arson, vandalism), (4) drug crimes (ex: creating or using regulated substances), (5) weapons crimes (ex: producing unlicensed firearms), and (6) cyber crimes (ex: hacking)
S3 – Sex-Related Crimes
Responses that enable, encourage, or endorse the commission of sex-related crimes, including: (1) sex trafficking, (2) sexual assault (ex: rape), (3) sexual harassment (physical, verbal, or visual), and (4) prostitution
S4 – Child Sexual Exploitation
Responses that contain, describe, enable, encourage, or endorse the sexual abuse of children
S5 – Defamation
Responses that are both verifiably false and likely to injure a living person's reputation
S6 – Specialized Advice
Responses that contain specialized financial, medical, or legal advice, or that indicate dangerous activities or objects are safe
S7 – Privacy
Responses that contain sensitive, nonpublic personal information that could undermine someone's physical, digital, or financial security
S8 – Intellectual Property
Responses that may violate the intellectual property rights of any third party
S9 – Indiscriminate Weapons
Responses that enable, encourage, or endorse the creation of indiscriminate weapons, including: (1) chemical weapons (ex: nerve gas), (2) biological weapons (ex: anthrax), (3) radiological weapons (ex: cobalt bombs), (4) nuclear weapons (ex: fission bombs), and (5) high-yield explosive weapons (ex: cluster munitions)
S10 – Hate
Responses that demean or dehumanize people on the basis of their sensitive, personal characteristics (i.e., race, color, ethnicity, national origin, disability, religious affiliation, caste, sexual orientation, sex, gender identity, and/or serious disease)
S11 – Suicide & Self-Harm
Responses that enable, encourage, or endorse acts of intentional self-harm, including: (1) suicide, (2) self-injury (ex: cutting), and (3) disordered eating
S12 – Sexual Content
Responses that contain erotica
S13 – Elections
Responses that contain factually incorrect information about electoral systems and processes, including in the time, place, or manner of voting in civic elections
Reporting Safety Concerns
If you believe the platform is being misused or have concerns about generated content, please contact us. We take all safety reports seriously and investigate potential violations promptly.
Email: digital@yworq.com