Next in thread → Next in month →

Re: [cti] Database Subcommittee / conceptual/logical model subcommittee

From
Jordan, Bret <>
Date
2015-07-07T17:36:40+00:00
ID
Thread
Re: [cti] Database Subcommittee / conceptual/logical model subcommittee
Great thoughts and ideas.  I just want to make sure the work gets done..  We need to find ways of greasing the gears, per say, of this effort.  The more HowTos we have to get people from zero to full speed the better.  

I understand the concern of building a reference implementation for say MongoDB and then everyone believing that, is the only way you can go.  However, I also see the other side of the coin, and a small shop, startups, and open-source developers are not going to really understand what is needed on the backend to store all of this data.  It seems like we need to help them figure that out some how. 

I think the UML models will greatly help this effort, but some best practices for a relational database and a document database would be a good idea..  Also some ideas of things to watch out for and reasons why you might choose one over the other.

Thanks,

Bret

Bret Jordan CISSP
Director of Security Architecture and Standards | Office of the CTO

Blue Coat Systems

PGP Fingerprint: 63B4 FC53 680A 6B7D 1447  F2C0 74F8 ACAE 7415 0050

"Without cryptography vihv vivc ce xhrnrw, however, the only thing that can not be unscrambled is an egg." 

On Jul 7, 2015, at 10:59, Barnum, Sean D. <> wrote:

I would tend to agree with Eric here on the risks of trying to specifyactual detailed implementations (e.g., a SQL schema that could be directlyinstantiated). As Eric points out, it could easily be assumed to be ³the²way you are supposed to do it rather than ³a possible² way to do it. It isalso likely to be tied to specific environmental/configuration assumptionsthat may not hold true across all SQL environments. As such it wouldlikely be somewhat brittle and difficult to maintain.My thinking, and what I plan to propose as part of the CTI STIX SCroadmap, is an approach where the SCs could take on work product efforts(if desired) to specify binding specs for particular database approaches(e.g. SQL, No-SQL document-centric, etc.) similar to the sorts of bindingspecs that will be created for particular schematic implementations (XML,JSON, etc.). These binding specs do not define the language but ratherprovide details of how you should implement various characteristics of thelanguage (defined in the language specs) in the specific technology beingbound to. For example, an XML binding spec for STIX would need to specifyhow to implement the Vocabularies data model (defined in the languagespecs) using XML Schema. It would also need to specify how to implementthe Controlled_Structure for data markings in XML. I think similar sortsof constructive guidance could be provided for best practices (hopefullyfrom real world lessons learned) for structuring the language in variousdatabase technologies. I think this would be very useful as the sort ofheadstart Bret is looking for (though certainly not turnkey) but wouldalso remain more flexible and easier to maintain across various specificenvironments or for minor language spec changes.What do you think?seanOn 7/7/15, 7:54 AM, "Eric Burger" <> wrote:I do not see this as in scope for CTI. This is precisely the kind ofthing we do not want to specify. The data format and protocol is fixed,so that if you use MySQL for your backend, I use NoSQL for my backend,and someone else uses 1,000 Bletchley Park Œcomputers¹ with note pads fortheir backend, we can still exchange information. Innovation occurs underand behind STIX/CybOX.I would offer that a *few*, independent, open source implementationswould be helpful to jumpstarting the ecosystem. A lot of the success toSMTP, SNMP, and IPFIX is because there were lots of freely availableimplementations for people to pickup and play with.Creating database schemas directly means making attendant technologychoices. To people outside the CTI group, it will look like these are the*only* acceptable technology choices. No matter how much we say ³theseare examples only, we expect you to innovate on your own,² it will nothappen. The same thing happens when we include sample code or protocolsnippets in the specifications. No matter how wrong the samples are andno matter how explicit we are in the specification that implementors needto follow the specification and not the samples, people will take thesnippets as gospel and not bother to read the specification.I am not sure if OASIS has this rule, but the IETF has a rule that aprotocol cannot become an Intenet Standard unless there are two,independent, interoperable implementations. TIFF never progressed toInternet Standard because everyone used Adobe¹s free libraries, and assuch there was never a second independent implementation. I fear that ifwe have an official data base schema, that will become *the* data baseschema. That would not be good for the ecosystem.All of that said, one of the things that made SIP successful is *outsideof the IETF*, although mentored with a lot of IETF folks, was a SIPImplementors group. That was where newbies could go to ask for basicadvice (³How do I parse XML?² ³How do I spell STIX?² ³What is a database?²) as well as more intricate questions (³I read the specificationand expected Œfoo¹ but I got Œbar¹ instead. What do you think I didwrong?²).The separation of SIP Implementors (not IETF) and SIP (IETF) was drivenby two factors. The first one was the SIP list became overwhelmed withquestions like ³Should I use C or C++ to build a stack?² That has nothingto do with protocol development. The signal to noise ratio fellprecipitously, and creating a SIP Implementors group significantlyimproved the S/N ratio for the people building, correcting, and adding tothe SIP spec. The second factor is the inverse S/N issue forimplementors. If you are a development manager who wants to know whatdata base to run on the back end of your aggregation system, you probablycould not care less if we are modeling a threat actor¹s left hand¹sacceleration as an entity or an attribute of a relation. You alsoprobably do not care if we are looking at the impact of delay tolerantnetworks for the interplanetary Internet. That is something important for*us* to think about as we get closer to a mission to Mars, but is mostlikely not important to someone who wants to deliver product this year.I have no problem with there being a CTI Implementors group. At this veryearly stage in CTI¹s development, I would offer it could be beneficialfor OASIS CTI to host the implementors group as a subcommittee. Thereason is we, as the protocol¹s developers, are learning a lot from earlyimplementations. However, I would also offer we spin it out as soon aspractical, so that the guidance provided is NOT interpreted as gospel.On Jul 6, 2015, at 5:49 PM, Jordan, Bret <>wrote:Since I proposed the idea of this working group 12+ months ago, andbegged Eric to run with it, a lot of what I was originally wanting andasking for has now been lost in really weird discussions about theobject model.  So lets rewind 12 months and get back to what I was asking for in thefirst place...  What I want out of this group is some guidelines and database schemasfor developers wanting to write TAXII / STIX / CybOX implementations.Basically, from a backend database standpoint, how do they get started?Which database systems should they use based on various implementationstrategies? What should the base configurations be for said databases?And maybe even some example implementations.For example, say an open source APP developer was going to write abasic STIX/Cybox Indicator/Observable UI that could read data in from aTAXII server, add comments and context and spit it back out, what typeof database should he/she use to store the data, and what should it looklike.For very simple things, where you are NOT doing all of STIX and Cybox,maybe a relational database would work fine and infant work really well.So in these cases, it would be nice if there was a .sql file thedeveloper could pull down that would build all of the necessary tablesstructure for him/her.  If they are doing something more complex, maybethey need a document database.  Then it would be really great if weproduced some documents / papers / implementations guides / and maybeeven some examples that could help them get up and running faster.The problems I see are:1) STIX is massive and very complex.  Just trying to learn it andfigure out how much of it you have to implement is a monumental task.2) Then if you need something in say Objective-C, you have to write theAPI to support STIX3) Then yo need to write a basic TAXII system to do what you want.4) Then you need to figure out what kind of database you are going touse to store the data.If we could help out on #4, that might just help make things easier forpeople getting started..Thanks,BretBret Jordan CISSPDirector of Security Architecture and Standards | Office of the CTOBlue Coat SystemsPGP Fingerprint: 63B4 FC53 680A 6B7D 1447  F2C0 74F8 ACAE 7415 0050"Without cryptography vihv vivc ce xhrnrw, however, the only thing thatcan not be unscrambled is an egg."On Jul 6, 2015, at 15:33, Barnum, Sean D. <> wrote:Eric, I definitely agree that there will need to be considerablecoordination between STIX and CybOX efforts. I would hope that that isnot a surprise to anyone. :-)seanFrom: Eric Burger <>Date: Monday, July 6, 2015 at 5:30 PMTo: "" <>Subject: Re: [cti] Database Subcommittee / conceptual/logical modelsubcommitteeEven if it means I¹m writing myself out of a Œchair,¹ I agree withSean. The most important task that people are talking about for the³database subcommittee² is the formal modeling of STIX/CybOX (and to alesser extent, TAXII). If that has to be a part of STIX or CybOX SCs,so be it.The downside is based on the work we have done so far at Georgetown,it is very difficult to build a model considering STIX and CybOX asseparate entities.On Jul 6, 2015, at 5:13 PM, Barnum, Sean D. <> wrote:We agree that these sub-topics have value and should be managedappropriately to ensure they are addressed consistently with minimalimpact to other areas and sub-topics.As the co-chairs for the CTI STIX SC we wanted to express ourthoughts on how these various sub-topics might be addressed mosteffectively.We propose that issues relevant to specific languages (languagespecifications at all levels of abstraction (ontology, logicaldata-model, specific schematic implementations (xml, json, etc.)),database and implementation guidance, etc.) should be managed fromwithin the appropriate STIX SC or CybOX SC. The concept of thelanguage specifications existing at different levels of abstraction isalready part of our way of doing things. For the last year or so wehave been working to lift the STIX specification from just XML Schemato a set of implementation-independent specifications based on a UMLmodel and explanatory textual documents. These specification documentswill form the normative basis of the language specification that willbe transferred to OASIS. It has also been the goal to eventuallyevolve the implementation-independent specification and lift it to amore formal and explicit semantic form. All of these differing levelsof abstraction are part of the effort to specify the language. Forconsistency sake they should all be part of a single evolutionarythread for the language and not separate parallel efforts. Similarly,specific guidance on database approaches or other implementationswould practically be tied to the language they are implemented tosupport and as such likely should fall within the scope of the SCworking on those languages. Within the language SCs these topics canbe broken down and managed using different work products asappropriate.Issues that tend to be relevant to the broader ecosystem (engagement,interoperability, etc.) may best be managed as separate SCs under theTC.We believe this approach will yield the best balance between focusingon specific issues, ensuring the right people are involved in theright efforts and achieving consistency across efforts and at thistime will likely improve our focus and support more rapid progress. Ifat some future time the TC decides that a different approach isneeded, it will be possible to modify the approach at that time.The first CTI STIX SC meeting next week will likely flesh out in abit more detail how we see this approach taking form for the STIX SC.We appreciate your consideration of our thoughts on the matter.Sean Barnum and Aharon CherninCTI STIX SC Co-chairsFrom: Patrick Maroney <>Date: Monday, July 6, 2015 at 1:03 AMTo: Eric Burger <>,"" <>Subject: Re: [cti] Database Subcommittee / conceptual/logical modelsubcommittee[+1]   "I have not been comfortable with calling this group the³database subcommittee² specifically because it is the data model, notthe data model implementation, that needs focus."[+1]  "...once people start looking at terms and concepts from amodel perspective instead of XML (or SQL, etc) data structures theydiscover issues, complexities, simplifications and opportunities thatare not very apparent looking at schema. This activity represents adifferent viewpoint that when combined with the more ³bottom up²implementation and representation concerns makes the specificationthat much better. For this reason it would be my suggestion that sucha viewpoint should drive the vocabulary and semantics and work inconcert with but not be the same as the team that focuses on the bestrepresentation and implementation in XML or a DBMS. "Patrick MaroneyOffice: (856)983-0001Cell: (609): <> on behalf ofEric Burger <>Sent: Sunday, July 5, 2015 3:17:55 AMTo: : Re: [cti] Database Subcommittee / conceptual/logical modelsubcommitteeI have not been comfortable with calling this group the ³databasesubcommittee² specifically because it is the data model, not the datamodel implementation, that needs focus. Cory nails it in one (secondparagraph below). In order to build real data migration tools, youreally need to understand what you are migrating. I would offer thefirst task (as opposed to a parallel sub-subcommittee) is to do themodeling.That is why we have been working on an OWL model for STIX/CybOX atGeorgetown. Our purpose was for a different goal, but the result couldbe generally useful.On Jun 24, 2015, at 4:45 PM, Cory Casanave <>wrote:Team,I purposely did not suggest a particular language for expressing theconceptual/logical model as that is a worthy topic of discussion forthe group. In the related OMG activity we are using a profile of UMLthat adds more semantic capabilities but has the tooling, establishedbase and graphic support of UML. This profile is currently goingthrough the standards process and is then able to generate OWL. Youcan say 90% of what you can say in OWL with less complexity. We havealso used OWL for other projects as it also has some valuablefeatures, but is also far from perfect.  This is a good topic fordiscussion. But, we get ahead of ourselves, the purpose and scopeshould drive such choices.What I have found in every similar activity is that once peoplestart looking at terms and concepts from a model perspective insteadof XML (or SQL, etc) data structures they discover issues,complexities, simplifications and opportunities that are not veryapparent looking at schema. This activity represents a differentviewpoint that when combined with the more ³bottom up² implementationand representation concerns makes the specification that much better.For this reason it would be my suggestion that such a viewpointshould drive the vocabulary and semantics and work in concert withbut not be the same as the team that focuses on the bestrepresentation and implementation in XML or a DBMS.In the best scenario the former would then generate the latter basedon transformation rules that map the terms, structure and semanticsonto the technology framework of choice. The existing schema providea valuable resource to start with whereas the models provide a betterway to evolve and certainly a better way to support multipletechnologies. This then can be considered a candidate strategy forphase 2, it is a different SDLC than starting with XML schema. Comingto consensus on our approach and SDLC should, perhaps, precedeforming subcommittees to start the work.Regards,Cory CasanaveFrom:  [mailto:] OnBehalf Of Jane Ginn - : Wednesday, June 24, 2015 2:05 PMTo: ; Jerome Athias; Cory CasanaveCc: ; : Re: [cti] Database Subcommittee / conceptual/logical modelsubcommitteeAll:Building on Cory's suggestion... Jerome's observations... and Sean'snote about using OWL or RDFS....Would it make sense to establish a Sub-Committee that combines someof the issues associated with database design that have beendiscussed previously (RDBMS vs. NoSQL) with this need forclarification at the abstract level (conceptual & logical)?If so.... would the scope of such a Sub-Committee also coverimplementation and tooling issues as was earlier suggested by Patrick?Further, what would be the tangible outputs, and how would they mapto the STIX/TAXII/ & CYBOX Sub-Committees?Jane Ginn, MSIA, MRPCyber Threat Intelligence Network, -------- Original Message --------From: "Barnum, Sean D." <>Sent: Wednesday, June 24, 2015 10:41 AMTo: Jerome Athias <>,Cory Casanave<>Subject: Re: [cti] Database Subcommittee / conceptual/logical modelsubcommitteeCC: Eric Burger<>," "<>I just wanted to add a note of clarification here for theintent/scope of STIX and CybOX to date.STIX and CybOX are intended to be Languages for expressing cyberthreat information and cyber observable information respectively.As such, they are more than simple data models or schemas. They alsoinvolve the conceptual model for their scope.To date, the emergent and exploratory nature of this communityseeking not only to formalize expressive representations for cyberthreat information but to work collaboratively and iteratively toeven figure out what that meant led to some necessary choices to workfrom the bottom up.This is why the language has initially been developed, refined anddefined in the form of XML schema. The schematic level of abstractiongave us something concrete to discuss, model specific technicaldetails and to experiment with real world data and implementations inorder to iterate and improve. XML schema was chosen not because it issome magical answer that everyone everywhere should use but ratherbecause it is ubiquitous, supported by a mature body of tooling andsynergistic standards (XPATH, Xpointer, Xquery, etc.) and provides apowerful formal schema language to explicitly constrain syntax whileenabling necessary flexibility. All of these things were needed tomodel and evolve a representation of an emergent knowledge spaceamong a very diverse set of players.This approach served us well to successfully get us where we aretoday but it has always been recognized that specifying the languageat this level of abstraction has significant downsides. First, it isdifficult to define semantics and high level concepts effectively atthis level and choosing any particular technical implementation (XML,JSON, etc.) inherently introduces technology-specific characteristicsthat really are not part of the more generalized language.In recognition of this, it has always been the plan to move thespecification of the languages to a more general form once anappropriate level of maturity and stability had been reached (verysimilar to the plan to move to a formal standards body at theappropriate time). The first steps toward this were put into motionseveral months back when work began on an implementation independentspecification for STIX and a separate but related one for CybOX. Itwas decided that based on community needs and maturity theappropriate first step in generalization would be to capture languagestructure and syntax in the form of a UML model that would beaccompanied by a set of textual specifications to explain andcharacterize the UML model in a more human consumable form. The draftset of these specifications for STIX 1.1.1 are currently available inthe STIXProject on github and the updated versions to STIX 1.2 shouldbe completed within the next couple weeks. This will be the primarynormative contribution to the CTI TC. There is a UML model for CybOXalso available but the set of accompanying full textual specs similarto STIX will not be created before transition to the CTI TC so thatwork will likely fall to the CybOX SC.While UML models are formal and are abstracted from particularsyntactic implementations (XML, JSON, etc.), they are not in allhonesty really built to convey high-level conceptual models orexplicit semantics of knowledge. They can be somewhat twisted toserve this purpose (as we have done in the implementation independentspecs) but the fact that they were designed to serve a systemsengineering rather than knowledge engineering purpose leads to someshortcomings. The inability of UML models to effectively conveyhigh-level conceptual models and explicit knowledge semantics in aformal fashion is one of the key reasons the textual specificationdocuments are required in addition to the UML. They not only providemore human-consumable characterizations of what is in the UML butthey are also needed to explain semantics that cannot effectively beexpressed in the UML. The upside is that some of these semantics cannow be explicit in the documents but it is in an informal form andstill open to human interpretation. What is ultimately needed for thelanguage specs is a way to formally express the full range oflanguage semantics and structure.I have personally asserted for a long while, and I know many in thecommunity agree, that the long term solution for specifications ofthe languages is to define and express them using mechanisms purposebuilt to define languages like this. That is, utilizing semanticforms of specification such as OWL and RDFS. These forms while lessfamiliar to many (part of the reason we decided to work from thebottom up) provide a way to clearly, explicitly, unambiguously andformally specify the high-level conceptual model for the languages,directly map it to any number of more detailed conceptual models, andthen directly map it to specific syntactic/schematic representations(logical models).Many members of the community have been eager to begin working atthis level but it was deemed important to first complete theabstraction work to the UML/textual specification level to serve as aXML-bias-free basis for initial semantic modeling. I propose thatsome of the CTI TCs early work should be focused on these activities.In fact, I would fairly strongly assert that many of the refactoringissues on the table for STIX 2.0 (e.g., abstraction of severalembedded structures (relationships, sources, assets, victims, etc.)to separate constructs) will require semantic modeling in order tofully understand and get right. I think the semantic discussions andmodeling as part of these activities could serve as some greatinitial steps towards more formal specifications for the languagesthat serve not only better integration for each language acrossabstraction levels (conceptual to logical) but also better alignmentand integration with related information representations within thecyber security sphere (MAEC, CAPEC, CVRF, OVAL, OpenIOC, etc.) andoutside the cyber security sphere.So, that was a long contextual way of saying that I strongly agreewith the need to understand and specify these languages across theabstraction spectrum (conceptual to lexical) but strongly feel thatthis should/must be done within the context of each language (I.e.within the STIX and CybOX SCs with cross coordination via the TC)rather than as a separate activity.SeanOn 6/24/15, 11:39 AM, "Jerome Athias" <> wrote:I'm a great fan of conceptual models!I skipped this step while reading the specifications to go directlytoa data relational model, but I can see a lot of benefits producing aCMap, especially for new adopters (just because one picture can tellthousands of words). It's easy to share also (e.g. CmapTools)The issue that I think we would encounter, is not so much about thelevel of abstraction (multiple CMaps could resolve that), whilethereis not so much concepts there (in CTI). (I used to do CMap forcomplexsystems)It is mainly, AGAIN, related to the taxonomy.You could see that when dealing with the extensions points, figuringout what would be the most appropriate standard/specification to mapCTI to. Things that are around CTI and that you have to deal with,such like Assets, Vulnerabilities, Exploits, Shellcodes, etc.But I assure you that it's fascinating ;)And while all these things are somehow linked together, it makesquitedifficult to make choice to -split- this into multiple models.(you could look at it in many ways, like asset-centric, risk-based,vulnerability-based, etc.)My 2c2015-06-24 18:18 GMT+03:00 Cory Casanave <>:There is certainly a value in a DBMS capability, perhaps one thatcan be implemented across multiple technologies. This may then alsorelate to the "conceptual model" initiatives which have alreadystarted. A conceptual model can bridge the exchange and repositoryviewpoints and also allow for greater flexibility in implementationtechnologies. We have had great success in generating schema aswell as transformations between them from models.With this in mind perhaps a conceptual and/or logical modelsubcommittee should be considered. Depending on the approach thiscould provide some of the value that is being sought for thedatabase. A separation of concerns would allow for the definitionof the database in models with implementation in one or more chosentechnologies. Such implementation would probably be anotheractivity.There is some grey area in what people call conceptual and logicalmodels and the levels of abstraction each represents. For me (andmany others), a conceptual model is a model of how the world isunderstood - it is then a model of the terms and concepts of theworld, not a data model. An "instance" of a person in a conceptualmodel is a real person - not data. A logical model is then atechnology independent data model about the world where choices aremade as to structure and representation. An "instance" of a personin a logical model is data. An initial activity of aconceptual/logical model subcommittee could be to define thepurpose, scope and appropriate level of abstraction.Of course the model activity is just as relevant to the exchangeschema and can help make them more understandable as well asprovide a basis for support of other technologies (essentially amodel driven architecture approach).  This works best when themodels are the normative definition and technology schema aregenerated from them. Since this tends to introduce more change (aswell as more consistency), it would best be coupled with the secondphase.There has already been work on conceptual models this directionseems consistent with the communities direction. With the above inmind we may want to consider a conceptual and/or logical modelsubcommittee.Regards,Cory CasanaveRepresenting OMG-----Original Message-----From:  [mailto:]On Behalf Of Jerome AthiasSent: Wednesday, June 24, 2015 7:06 AMTo: Eric BurgerCc: : Re: [cti] Database SubcommitteeI wonder if providing consumer-oriented XQuery examples (maybewith the STIX idioms) would help providing guidance andtest/validation cases2015-06-22 14:20 GMT+03:00 Eric Burger<>:Jerome (as he often does) gets this right in one (how about that- use a British colloquialism instead of a US one!).We just submitted a paper for publication at MILCOM looking atSTIX/TAXII/CybOX versus IODEF/RID from the perspective of humansversus machines doing the processing. My guess is you can guessthe end of the story: STIX/TAXII/CybOX is much better formachines. IODEF/RID is much better for people. Since the goal isfor inter-machine communication, you get the point.It does mean there is a lot riding on VERY clear, implementable,interoperable specifications. Debugging this stuff is going to bea nightmare, more especially if the language is so nuanced thereare dozens of ways of saying the same thing.---------------------------------------------------------------------To unsubscribe from this mail list, you must leave the OASIS TCthat generates this mail.  Follow this link to all your TCs inOASIS at:https://www.oasis-open.org/apps/org/workgroup/portal/my_workgroups.php---------------------------------------------------------------------To unsubscribe from this mail list, you must leave the OASIS TC thatgenerates this mail.  Follow this link to all your TCs in OASIS at:https://www.oasis-open.org/apps/org/workgroup/portal/my_workgroups.php---------------------------------------------------------------------To unsubscribe from this mail list, you must leave the OASIS TC that generates this mail.  Follow this link to all your TCs in OASIS at:https://www.oasis-open.org/apps/org/workgroup/portal/my_workgroups.php 

Attachment:
signature.asc

Description: Message signed with OpenPGP using GPGMail
Next in thread → Next in month →