Why so many ‘fake’ Big Data Gurus?
Where do you all come from?
Where do you all come from?
All your integrity’s gone
Now tell me, where do you all come from?
From ‘Where Do You All Come From‘ by Mott the Hoople (more…)

Why so many ‘fake’ Big Data Gurus?
Where do you all come from?
Where do you all come from?
All your integrity’s gone
Now tell me, where do you all come from?
From ‘Where Do You All Come From‘ by Mott the Hoople (more…)

If you enjoy this piece or find it useful then please consider joining The Big Data Contrarians:
Join The Big Data Contrarians here: https://www.linkedin.com/grp/home?gid=8338976
Many thanks.
Now everyone is doing Big Data you don’t want to be the odd one out, right? Of course not.
Now, if you are serious about looking at Big Data from a business perspective then I will try and lend you some advice. If you are doing it from an IT or technology perspective, then I wish you good luck, and I hope that your Big Data initiative doesn’t turn into another tech crash-and-burn show.
Now some Big Data pros are telling us that the place to start with Big Data is with strategy. Now, I’m too polite to call this out as abject bullshit, even though it is, and will instead content myself by offering an alternative and simple approach to approaching and addressing Big Data.
My first piece of advice is this. DON’T START WITH STRATEGY!
Strategy is a coherent, cohesive and executable response to a significant challenge.
Strategy is not a definition of objective, a wish list of what you are trying to achieve or aspirational goals of a nebulous nature. No, strategy is not the objective but a means of reaching that objective. Strategy is real, tangible and executable. Strategy is doing.
So what is a Big Data strategy?
If a company is looking at the Big Data options, the last place they should want to start out from is from strategy. That is as silly idea as they come. Starting with strategy on the road to formulating viable responses to significant challenges and opportunities is like saying that before we choose strategic options and a realisable strategy, then you must have a strategy in place.
Strategy is not working out what you want to achieve. That sort of thing should happen prior to any strategic work. Neither is strategy an exercise in establishing starting points, nor formulating questions nor understanding the challenges. All of this should come well before the major strategy aspects even kicks-in.
Big Data strategy is a realisable, tangible and manageable response to a significant challenge, one that depends heavily on the availability, usability and credibility of Big Data (or Very Large Data Bases) and the business value of processing that Big Data.
So, a word of advice. If you are thinking of embarking on a Big Data initiative, do not start with strategy. That is a really daft place to start.
Start here instead. With real business imperatives. This is where you are thinking about the big and significant challenges to the business, and how, at a high level of abstraction, you could go about meeting those challenges. Here you identify your challenges and your responses, aligned to your objectives.
If you can identify business imperatives that make it absolutely necessary to include elements of Big Data, then go forward with that mandatory requirement in mind. If not, then don’t try to shoe-horn Big Data into a place where it really isn’t needed or wanted. Because if you go against the grain in this way it may well hurt you and your business, in more ways than you bargained for.
In order to go out looking for data requirements driven by business imperatives, we really need to know what we are looking for.
What we are looking for maybe highly tangible or less so. We may have to derive the data we are looking for by refining, aggregating, enriching, filtering and cleansing. Therefore, with those and other aspects in mind, we can go out and find what we need.
From looking at the data requirements, you should have a good idea of potential sources of that data. Agility in this aspect is predicated on the premise that one knows the systems on the IT landscape, the business processes and all the potential sources of data – at a high level at least. So, this is not the sort of work you can do remotely with little or no knowledge of the clients business, IT setup, processes or culture.
But anyway, after you identify the sources you move on to the next step.
Here you discuss aspects of the data you require with the database / application platform owners to ensure that:
So far so good. Once passed these hurdles (and don’t forget this is a super-simplification) we move in to the next.
So, now we know:
Therefore, we go ahead and create a proof of concept or three. Simples!
However, make sure that all prototypes are governed by these simple timeless guidelines:
You run your proof of concept. You analyse, assess and represent your outcomes. You socialise, present and interpret.
When you’ve done that you are in now in a good position to estimate the usefulness of the exercise, from both a qualitative and quantitative perspective.
I did not want to touch in specific aspects of technology in this piece, in part, because I did not consider it a central issue in the theme of things. Of course, as part of creating proofs of concepts and pilot schemes you may want to experiment with the swatch (swaith? oh for auto-correction) of technologies out there. So go ahead and evaluate ‘Big Data’ technologies, and don’t forget, the answer to every Big Data technology question isn’t an automatic ‘Hadoop’. There are other valid Big Data technology options around, such as Lustre and GPFS, or even Oracle, Teradata or EXASol. Also, remember this, if all you are working on is a prototype, a proof of concept or a pilot then you can try and negotiate a free license with any of the major DBMS vendors for that initiative. So negotiate, bargain and get the most appropriate technologies with the best deals.
Finally I will leave you with three guidelines to consider:
Many thanks for reading.
In subsequent blog pieces I will be sharing my views on the evolution of information management in general, and the incorporation novel and innovative techniques, technologies and methods into well architected mainstream information supply frameworks, for primarily strategic and tactical objectives.
As always, please reach out and share your questions, views and criticisms on this piece using the comment box below. I frequently write about strategy, organisational, leadership and information technology topics, trends and tendencies. You are more than welcome to keep up with my posts by clicking the ‘Follow’ link and perhaps you will even consider sending me a LinkedIn invite if you feel our data interests coincide. Also feel free to connect via Twitter, Facebook and the Cambriano Energy website.
For more on this and other topics, check out some of my other posts:
Absolutely Fabulous Big Data Roles – https://www.linkedin.com/pulse/absolutely-fabulous-big-data-roles-martyn-jones?trk=prof-post
Not banking on Big Data? – https://www.linkedin.com/pulse/banking-big-data-martyn-jones?trk=prof-post
10 amazing reasons to join The Big Data Contrarians –https://www.linkedin.com/pulse/10-amazing-reasons-join-big-data-contrarians-martyn-jones?trk=prof-post
Amazing Data Warehousing with Hadoop and Big Data –https://www.linkedin.com/pulse/cloudera-kimball-dw-building-disinformation-factory-martyn-jones?trk=prof-post
The Big Data Contrarians: The Agora for Big Data dialogue –https://www.linkedin.com/pulse/big-data-contrarians-agora-dialogue-martyn-jones?trk=mp-reader-card
The Big Data Shell Game – https://www.linkedin.com/pulse/big-data-shell-game-martyn-jones?trk=mp-reader-card
Aligning Data Warehousing and Big Data –https://www.linkedin.com/pulse/aligning-data-warehousing-big-martyn-jones?trk=mp-reader-card
Big Data Luddites – https://www.linkedin.com/pulse/big-data-luddites-martyn-jones?trk=mp-reader-card
Data Warehousing Explained to Big Data Friends –https://www.linkedin.com/pulse/data-warehousing-explained-big-friends-martyn-jones?trk=mp-reader-card
Big Data, a promised land where the Big Bucks grow –https://www.linkedin.com/pulse/big-data-promised-land-where-bucks-grow-martyn-jones-6023459994031177728?trk=mp-reader-card
The Big Data Contrarians – https://www.linkedin.com/pulse/big-data-contrarians-martyn-jones?trk=mp-reader-card
Is big data really for you? Things to consider before diving in –https://www.linkedin.com/pulse/big-data-really-you-things-consider-before-diving-martyn-jones?trk=mp-reader-card
Big Data Explained to My Grandchildren – https://www.linkedin.com/pulse/big-data-explained-my-grandchildren-martyn-jones?trk=mp-reader-card

If you enjoy this piece or find it useful then please consider joining The Big Data Contrarians:
Join The Big Data Contrarians here: https://www.linkedin.com/grp/home?gid=8338976
Many thanks.

Plus ça change, plus c’est la même chose.
Jean-Baptiste Alphonse Karr
I wrote a piece called ‘7 New Big Data Roles for 2015’. I published it on LinkedIn. Many people read it. Some people made suggestions. Others politely ignored it.
I listened to the suggestions, comment and criticisms, and revised the piece as a result.
So here, it is… I hope you like it. And if not, I might try again in six months’ time.
(more…)
Listen up Big Data playmates! The ubiquitous Big Data gurus, tied up in their regular chores of astroturfing mega-volumes, velocities and varieties of superficial flim flam, may not have noticed this, but, Hadoop is getting set up for one mighty fall – or a fast-tracked and vertiginous black run descent. Why do I say that? Well, let’s check the market. (more…)

If you know all about Sentiment Analysis, you’ve come to the right place. Because I don’t have a clue if what I know about it is accurate or not.
I started to do a bit research into this Sentiment Analysis lark, in particular with the theoretical idea of using it to analyse and draw conclusions from comments on Pulse – assuming that this is what it can be used for.
To begin at the beginning, which is good place to start, I read the piece on Wikipedia, and this was how it began:
“Sentiment analysis (also known as opinion mining) refers to the use of natural language processing, text analysis and computational linguistics to identify and extract subjective information in source materials.
Generally speaking, sentiment analysis aims to determine the attitude of a speaker or a writer with respect to some topic or the overall contextual polarity of a document. The attitude may be his or her judgment or evaluation (see appraisal theory), affective state (that is to say, the emotional state of the author when writing), or the intended emotional communication (that is to say, the emotional effect the author wishes to have on the reader).” Source: Wikipedia Link:http://en.wikipedia.org/wiki/Sentiment_analysis
Well, that’s a fairly intuitive description. I could have almost have guessed as much.
But, back to the aim of analysing sentiment in Pulse comments, where to start and what to do.
What would sentiment analysis make of these:
On the death of an IT-business celebrity. What would sentiment analysis make of the very emotive comments of desolation, sadness and poignancy of people who didn’t personally know the departed, even remotely, or maybe didn’t even know of them until after they had ‘shuffled off life’s mortal coil’? How would that work? What would sentiment analysis make of the maudlin aphorisms, surrogate grief and bizarre sorrow of people separated by more degrees than Kofi Anan and Mork from Ork. What additional insight does sentiment analysis tell us when these comments are analysed along with the body of the text and other comments that triggers these comments?
In a similar vein, how does sentiment analysis catch instances of sycophancy? Especially considering the fact that some of it is so ‘in your face’ and blatant that it often times seems to be a bad parody of a bad parody. “Oh, Ricky, why are you such a sexy brainbox?” How does it work in those situations?
Worse than that is the preening, gushing and obtuse texts of massive, errm… fabulators[i]. If it wasn’t about Big Data or Strategy or IT, it would be about something else, usually about the writer themselves. “I give Rafa and Rodge tips on tennis! I went to the University of the Universe and got a first! I challenged Superman to a race, and won! I have read the entire works of Dan Brown, 25 times…Neeeh!” What would sentiment analysis do with that sort of gold?
Also, what does sentiment analysis do with texts so ambiguously daft that they could mean anything? Okay, it might be able to pick up a few trigger words here or there, “rubbish”, “of”, “load”, “a”, “what”, etc. However, how does it know when “excellent” is being used in a way that means anything but excellent? For example, “Excellent Big Data job there”, with the silent “if you want a job doing properly then do it yourself”.
Finally, for the purpose of this little piece, what would sentiment analysis do with term abuse, if it could actually identify it? Going back to the use of the terms such as Big Data or Strategy, how can sentiment analysis discern between the dopey and wrong-headed use of the term, and when it is actually being used in a coherent, cohesive and consistent way, in line more or less with its formal definition? I suppose we can always write a mountain of rules to help us out:
If topic in focus of piece is strategy
And context of topic is business
And author of piece is Richard Rumelt
Then the credibility of text is good (with a certainty of 100%)
But you and try and maintain a rule base with isntances like that. It soon becomes a management nightmare.
Alternatively, maybe it could be used to analyse this text. It’ll have its work cut out, that’s for sure. Does sentiment analysis do sarcasm and cynicsm?
Anyway! I bet you might know how this sentiment analysis works, don’t you? On the other hand, if not, then it will be someone else who ‘knows’. But of course, all will not be revealed, because it’s a secret so powerful, that in the wrong hands it could be used to dominate the entire galaxy.
Only joking; and many thanks for reading.
[i]To engage in the composition of fables or stories, especially those featuring a strong element of fantasy: “a land which … had given itself up to dreaming, to fabulating, to tale-telling” (Lawrence Durrell).
lang: en_US

Most of us would probably like to work in a profession recognised for its legality, decency and honesty. At least I hope so. In my line of work, what we have right now is palpable evidence that the IT industry lacks a moral compass.
Imagine this. A major sensationalist tabloid pulls together a team of diverse journalists who are set to work on a national campaign to promote very high usage of sunbeds as a cure for cancer. Why? The newspaper owner’s son owns the sunbed franchise.
The health experts criticise the publisher for being irresponsible, unprofessional and lacking in scruples.
The public is mainly undecided, but many take the story on face value and adopt the fad. The intensive use of sunbeds sharply increases. Elsewhere, in unrelated news, the cases of skin cancer show a marked increase. Some blame it on EU legislation for bangers and bananas.
In spite of protests, the press campaign continues over many months.
Eventually, and based on the evidence of recognised health experts and bodies, the press regulatory association tries to get the offending publisher to temper their claims, but without any success. It is only when the government’s lawyers step in and threaten the newspaper owners with legal proceedings, do they freeze their campaign. Much later, the editor resigns and the board of directors issue a short apology on the back pages of their much vaunted organ.
We have that in IT. Our current sunbed cure for cancer, if you believe those who are ‘bigging it up’, is undoubtedly Big Data.
I occasionally post content to Linkedin, some of it (maybe even this piece) gets promoted through the Pulse Big Data channel. There are some reasonable pieces pinned to that channel, but unfortunately, for much of the time what we get is total and moronic Big Data astroturfing. Tantamount to the equivalent of Big Data’s very own Big Lie campaign.
The Linkedin Big Data channel reflects life, and it is full of self-aggrandising and shameless marketing guff, shot-through with scandalously flimsy promotions of tendentious success stories, specious claims of value, half-truths about realisable benefits and embarrassing conjecture about the importance of social media and internet logs.
What I am referring to mainly are superficially neutral (yet virally toxic) pieces placed in the public domain in order to promote Big Data at any cost.
Now let’s step back a bit.

For over 125 years, the Financial Times (FT) has built up a solid professional reputation for accurate reporting, reliable journalism and informative editorials. The FT is a newspaper trusted by its discerning readership and admired everywhere. In fact, I could not imagine their journalists writing about markets, securities and financial houses the same way that pundits elsewhere write about Big Data, Dark Data and the Internet of Things. Because the FT knows, that maintaining the trust of their readership is far more important than winning the short-term favours of a few market players.
So consider this; if we in IT cannot bring our standards of communicating with the public up to the levels of the financial industry, at minimum, you know what that means don’t you?
Exactly. The IT industry will have a far worse public image problem than the bankers and the solicitors currently have, and we all understand the general public appreciation of those professions.
Now, call me old fashioned, but for me that possibility is worthy of serious consideration, and especially by those in IT who confuse no holds barred pimping of fads, trends and technology, in which truth, decency and honesty are optional, for ethical, candid and informative analysis and reporting of the industry.
How will the industry take these criticisms?
To go back to the sunbed analogy what we will most certainly get comments in this vein:

Whilst those who rail against ‘the cancer curing advantages of sunbed use’ may be right – or at least partially right – the sunbed revolution will continue, just as the IT revolution industry has done, and in spite of people saying that the age of computing would be a passing mania.
So, when someone tells you “intensive sunbed use is just a dangerous fad”, what they actually mean to say is that we don’t need the term any more, as intensive sunbed use is here to stay, as are those who are shrewd, unprincipled and cynical enough to cash-in on the public’s gullibility and wilful stupidity when it comes to fads.
Yes, it does get that bad.
We have people who seemingly spend all their waking lives working out not-so-original ways and means of riddling the IT industry with vacuous bullshit, and what Big Data promotion has shown us clearly is that what we have palpable and comprehensive evidence that the IT industry in general lacks a moral compass.
Is that a reflection of IT, of those who create and manipulate IT fads, or of society in general?
Many thanks for reading.
As always, please share your questions, views and criticisms on this piece using the comment box below. I frequently write about strategy, organisational leadership and information technology topics, trends and tendencies. You are more than welcome to keep up with my posts by clicking the ‘Follow’ link and perhaps even send me aLinkedIn invite. Also feel free to connect via Twitter, Facebook and the Cambriano Energy website.


I have written at length about the fundamental contradictions of Big Data, but what I have omitted in the past is quite possibly the biggest contradiction of all. Probably because it has more to do with how Big Data is continually hyped, rather than having anything to do with Big Data as a bag of technologies – which has a whole assortment of problems in its own right.
Last time I spoke with you about the contradictions of the Big Data it was about the three Vs of volume, variety and velocity. In general, it was a view that was well received, even if not widely understood. Which of course is close enough for government work. But, get ready for “something completely different”.
If on the one hand some folk can claim that Big Data provides fact based insights and reliable forecasts of future habits, trends and preferences, then why is it so difficult to produce and socialise – yes, I like to use that term – Big Data success stories?
In short, I think we have arrived at the stage in Big Data’s cycle where it is reasonable to ask pundits to either put up or shut up.
So, why aren’t the current Big Data success fables accompanied by facts, such as names of those involved (at least businesses), the sponsors, the suppliers, the purpose of the exercise, the desired outcomes, the data used, how it is processed, what the results were, and what tangible benefits, if any, were accrued or are accruable. If that is not enough, then let people mention the technology used, the products purchased or licensed and the methodology followed.
In short, what I would like to know is why are the evangelists of Big Data telling us that bigger data is better, that more variety leads to greater insight, and that velocity is king. Why do we we told that Big Data almost assuredly results in better decisions, by people who are coy, shy or secretive about almost facts and data coming out of Big Data projects?
I have been reminded, time and time again, that there are Big Data success stories out there, and I have even been told that this information would be fully shared with me once it was agreed with the ‘clients’ that it was okay to do so. Okay, that’s fine, I know Big Data is a roaring success story (at least in people’s minds,) and I also know that it takes some time to make things up – some people are just not creative. Sure, I was told about these ‘successes’ some time ago, and you know, I’m not expecting anything that’s worth shaking a stick at, either now or later, but I’m still waiting, boys. Notwithstanding, you will still called you out as vacuous bullshitters when the time comes.
“But” I hear you cry “there is a wealth of success stories in the presses”.
Well, no, and you would wrong and gullible and foolish to think, but that is your problem, but unfortunately also mine, because this is my profession that you are playing fast and loose with.
The fact is that there is “wealth” of content that people try and pass off as legitimate Big Data success stories, but they aren’t in fact success stories, in any way, shape or form.
The thing is, people may read the blog title and even the stand-fast, but will be less inclined to actually read the article, so what remains is the impression that there are ‘loads of Big Data success stories’. But if people actually read the articles and were intelligent enough to understand them, then they would realise that inevitably there is a massive mismatch between the title of these pieces and the content. Indeed, if these pieces were actually pieces of advertising, rather than blog comments, they would be denounced in some jurisdictions for not fulfilling the advertising criteria of legal, decent and honest.
There is one more thing that Big Data evangelists (or any self-styled pundit, guru or expert for that matter) should understand, internalise and remember. If you say that you have a Big Data success story, with all the details, and that isn’t in fact the case, and it isn’t even remotely a success story or even true, then you are simply lying, and that’s deceit, it’s unprofessional and it’s unethical, and you are a scoundrel. So live with or fix it, the choice is yours.
Many thanks for reading.

Many people come up to me in the street and ask me what Big Data is all about. It has happened to me so many times in the past that I am convinced that it might just happen to you as well. I know sort of thing, I read the Big Data tealeaves. Nothing gets past me.
The first time a complete stranger came up to me in public and said “Hello, will you tell me what this Big Data lark is all about then?” I was lost for words, you just ask my Aunt Dolly, he can vouch for that, no problem. Later that day I read a book – it was my dad’s book – and I then decided to adopt a strategy.
Therefore, in the spirit of springtime goodwill to all men and women, I have put together this blog piece in that hope that it will enlighten, help and entertain.
What is big data?
Big Data can be characterised by the 10 Vs – yes, 10, not 4. Which, in my book, is more than enough to bring up-to-speed the average Big Data John or Jane that one meets on the street, and who naturally wish to be informed of such matters.
In layperson’s terms this a series of landmarks and pointers in the analytics space used to frame and guide the didactic aspects of Big Data.
The fundamental Vs of the Big Data canon are these:
So, let me now explain what each of these characteristics mean to those who might know and for those who might want to know.
Vagueness: This is perhaps the trickiest of questions to address, given the vast panorama that is cast before this incredibly complex yet easily graspable concept. But let me state this, and let there be no mistake about it. At this point in time, what makes Big Data vague is also what makes Big Data specific, explicit and certain. That is to say, in order to ‘come to an understanding’ of Big Data, it is necessary to completely embrace the dialectic of knowing the unknowable. So belief is an absolute essential element – belief and data, that is.
Volume – If there ever was a time to “pump up the volume”, we have it here with Big Data.
Big, voluminous, gorgeously rotund and infinite. Big Data is called Big Data because there is a lovely, roly-poly, likeable never-ending load of it. Its volumes can be measured in zeta-bytes, which you can be assured, is a helluva lot of data.
Variety – As they might say down my way, “variety is the spice of life, innit”. This is what makes Big Data so special. So appealing.
Because before Big Data there was absolutely no variety in anything, at all. We lived in a bland world, bereft of detail, nuance and diversity. Nothing could be measured, analysed or explained, because we lacked Big Data. We were ignorant. So ignorant and stupid that we couldn’t see the sense of putting the diapers next to the beer, or of offering three for the price of two.
Fortunately, today this is no longer the case if we don’t want it to be, and thanks to Big Data we have a veritable sensorial explosion. No longer is IT just a couple of symbols scribbled in crayon on someone’s school notebook.
Virility – Move over Smart Data, the new kid on the block is Big Data.
If Big Data were described in the manner of a religious text, it would be accompanied by a never ending narrative of begets.
So, what does that mean?
Simply stated, Big Data creates itself, in and of itself. The more Big Data you have, the more Big Data gets created. It’s like a self-fulfilling prophecy in 360 degree, high-definition, poly-faceted and all-encompassing knowing. The sort of thing that governments would pay an arm and a leg to get their mitts on.
Velocity – Velocity is of the essence. Velocity kills the competition. More velocity, less haste.
We demand that service is ‘velocious’. ‘Everything’ must be ‘now’, or it’s too late.
This means we need to be able to handle Big Data at velocity – at the speed of need.
Charles Babbage once stated (or maybe it was more than once) that “whenever the work is itself light, it becomes necessary, in order to economize time, to increase the velocity.”
But remember, we are dealing with mega-velocity here, so don’t drink and drive the Big Data Steamship, Star-ship or Mustang.
Vendible – If you can sell it, and sell it as Big Data, then it ‘is’ Big Data. If you can’t, then it’s not. The saleability of Big Data proves its existence.
So, what are the vendible aspects of Big Data?
Let’s leave that easy question for another day. But for now I can confidently state that it is used to mobilise armies of commentators, industry analysts, publicists, punters, writers, bloggers, gurus, futurologists, conference organisers, conference speakers, educators, customer relationship managers, salespeople, marketers and admen.
Vaticination – Edmund Burke is down on record as stating that “you can never plan the future by the past”. Now Burke may have been a clever person when it came to many things, but he wasn’t exactly a whiz when it came to Big Data.
There are people in the world who are in no doubt that Big Data provides the sort of visionary and predictive powers only previously obtainable through ritual sacrifice, magic potions and the casting of spells. Others are highly critical of the understatement implicit in this belief.
For many, Big Data will make the Oracle of Delphi look like a mere call centre.
This is why the power of vaticination plays a characteristically important role in the world of Big Data.
Voracity – This is based on the quasi-rationalist argument that Big Data is big and it has an omnipresent and insatiable self-fulfilling desire.
Big Data comes with an attendant requirement for hardware, even if it is a whole load of consumer hardware tacked together in a magnificent and miraculous mesh of magic.
Big Data can be characterised by voracity, but this comes hand in hand with the ‘ventripotent’ IT industry.
Veracity – The eminence of the data being captured for Big Data handling can vary significantly. The quality or lack of quality of the data naturally has the potential to impact the accuracy of analysis using that data.
Before Big Data arrived on the scene we knew nothing about Data Quality or data verification. This is why ETL and Data Cleansing tools lacked the power to effectively quality check and verify data, to ensure that any erroneous or anomalous data was rejected or flagged.
But now, with the sophistication of tools such as ‘grep’ and ‘awk’ at our disposal, we have the power in our hands to ensure nothing ‘dodgy’ gets into the analytical mix.
Vanity – In my opinion, to fully grasp the underlying and profound meaning of Big Data, it is essential for us to understand the difference between vanity and conceit. Max Counsell claimed that “Vanity is the flatterer of the soul”. Goethe characterised vanity as being “a desire for personal glory”. After an incident with an Anarchist (presumably a Big Data Anarchist), Blackadder remarked to Baldrick that “The criminal’s vanity always makes them make one tiny but fatal mistake. Theirs was to have their entire conspiracy printed and published in plain manuscript”.
So that ends the brief rundown of the defining characteristics of Big Data.
So, to summarise. That, which has passed before, necessarily divulges both the upside and downside of Big Data. By reaching out, opening up the kimono and relating the 10 Vs we are disclosing that which cannot be disclosed, exhibiting the absence of essential essence, and thereby opening up the entire field, discipline, profession, science and art to examination, questioning and ridicule.
Many thanks for reading.

Consider this. Why be shameful when all around you have apparently no idea of right from wrong?

You are the boss. You are the leader, coach and manager, and there are some things that you just got to learn, like it or not. One of these skills is to be able to identify when someone has quit. “How dare they?” I here you ask.
The first time I quit a job and didn’t tell anybody was when I was in the RAF working as a fighter pilot in World War 2, and I accidentally bombed Newport in South Wales, and was given a stern talking to for my troubles. Well, I didn’t actually quit and I was never in the armed forces and I was born into the era of the Beat Generation, but that’s by the by, it’s just there for effect, to create some artificial empathy between me and those who have actually quit a job and not told anyone about it. Myself, I would never do such a thing. Although to be fair, Newport has looked like it has been freshly bombed with dark green, brown and grey shades of poster paints and self-raising flour, since forever. (more…)

Dans ce pays-ci, il est bon de tuer de temps en temps un amiral pour encourager les autres – Voltair
My gran used to tell me that honesty pays. Of course, she never really understood banking or IT, probably because she didn’t want to know anything about them, and she never lived to witness the amazing hype circuses, the spin doctors spiel or the focus-group dog-and-pony show of the 21st century. Indeed, if honesty were a guaranteed payer my gran would have amassed more wealth than even Warren Buffet himself.
If my gran lived today, she might reflect on what Big Data might be about – maybe she would even consider it benignly, as a sort of shelter for fallen men of once uncertain virtue. We will never know. So onwards and upwards.
The Harvard Business Review contemplated honesty in somewhat different terms:
“Honesty is, in fact, primarily a moral choice. Businesspeople do tell themselves that, in the long run, they will do well by doing good. But there is little factual or logical basis for this conviction. Without values, without a basic preference for right over wrong, trust based on such self-delusion would crumble in the face of temptation.”
In a marvellous book, A few good from Univac, David E. Lundstrom narrates the story of Sperry Univac in the 1960s, one of the true great innovators in the first forty years of IT, and includes an allegory taken from the engineering front-line. I will recount it here, edited to highlight the zeitgeist, for your entertainment and as Voltaire put it, “to encourage the others”:
In the beginning was the Big Data Plan.
And then came the Big Data Assumptions.
And the Assumptions were without form.
And the Plan was without substance.
And darkness was upon the face of the Workers.
And they spoke amongst themselves, saying: “It is a crock of shit, and it stinketh.”
And the workers went unto their Supervisors and said: “It is a pail of dung, and none may abide the odor thereof.”
And the Supervisors went unto their Managers, saying: “It is a container of excrement, and it is very strong, such that none may abide by it.”
And the Managers went unto their Directors, saying: “It is a vessel of fertilizer, and none may abide its strength.”
And the Directors spoke amongst themselves, saying to one another: “It contains that which aids plant growth, and it is very powerful.”
And the Vice Presidents went unto the President, saying unto him: “This new plan will actively promote the growth and vigor of the company, with powerful effects.”
And the President looked upon the Big Data Plan, and saw that it was good.
“But?” I hear you say, “why fight it, why not take advantage of the Big Data zeitgeist?”, “Why not cash in on the grand bonanza Big Data bandwagon?” or “Monetise the 3 three famous Vs of Big Data?”
Well, it had crossed my mind, briefly, and (outside of the USA) we’ve all done stuff we have not entirely believed in, so the temptation to cash in is present, capisci? This paraphrasing of a piece from My Blue Heaven might give you a better idea:
One of my best friends makes his living as a completely phony Big Data Scientist. For two hundred bucks he can make you a Data Scientist or a Big Data guru. Some guys give you an education but this guy gives you immediate access to high paying jobs, sex that would make the 256 trillion Shades of Blah blush and a life in the City, the Big Apple or a small town in Germany.
Moreover, for an extra 250 bucks (limited time offer) you can also become a certified Big Data Neuro Trainer, which will allow you to do unto others what has been done unto you.
I also considered Big Data Brokerage, Big Data Certification and Big Data Independent Trading (New York – Paris – Peckham). The opportunities are immense.
However, what happens when the Big Data well runs dry, and I (and many others get tarnished with the mark of Big Data) become pariah by complicity, collusion or simple association?
That question I will leave for another day. But just consider the following.
All right, I admit, I am a big long-time fan of comic genius Mel Brooks, who has a knack of capturing deep insight from the human condition, especially when the human condition is off guard and shallow. In that vein, this is how I like to think the dialogue from the Dole Office scene from The History of the World Part Two would have gone, if he were to write that today:
Dole Office Clerk: Occupation?
Data Magnus Comicus: Stand-up Big Data scientist.
Dole Office Clerk: What?
Data Magnus Comicus: Stand-up Big Data scientist. I coalesce the vaporous datas of the human interaction with the social-media networking, Internet of Everything, and always-connected experience into a… viable, analytical and meaningful predictive-comprehension.
Dole Office Clerk: Oh, a Big Data bullshit artist!
Data Magnus Comicus: *Grumble*…
Dole Office Clerk: Did you bullshit Big Data last week?
Data Magnus Comicus: No.
Dole Office Clerk: Did you try to bullshit Big Data last week?
Data Magnus Comicus: Yes!
Finally, I leave you with some wise words from Israeli American professor of psychology and behavioural economics, Dan Ariely:
“Big data is like teenage sex: everyone talks about it, nobody really knows how to do it, everyone thinks everyone else is doing it, so everyone claims they are doing it…”
Many thanks for reading.

Dark data, what is it and why all the fuss?
First, I’ll give you the short answer. The right dark data, just like its brother right Big Data, can be monetised – honest, guv! There’s loadsa money to be made from dark data by ‘them that want to’, and as value propositions go, seriously, what could be more attractive?
Let’s take a look at the market.
Gartner defines dark data as “the information assets organizations collect, process and store during regular business activities, but generally fail to use for other purposes” (IT Glossary – Gartner)
Techopedia describes dark data as being data that is “found in log files and data archives stored within large enterprise class data storage locations. It includes all data objects and types that have yet to be analyzed for any business or competitive intelligence or aid in business decision making.” (Techopedia – Cory Jannsen)
Cory also wrote that “IDC, a research firm, stated that up to 90 percent of big data is dark data.”
In an interesting whitepaper from C2C Systems it was noted that “PST files and ZIP files account for nearly 90% of dark data by IDC Estimates.” and that dark data is “Very simply, all those bits and pieces of data floating around in your environment that aren’t fully accounted for:” (Dark Data, Dark Email – C2C Systems)
Elsewhere, Charles Fiori defined dark data as “data whose existence is either unknown to a firm, known but inaccessible, too costly to access or inaccessible because of compliance concerns.” (Shedding Light on Dark Data – Michael Shashoua)
Not quite the last insight, but in a piece published by Datameer, John Nicholson wrote that “Research firm IDC estimates that 90 percent of digital data is dark.” And went on to state that “This dark data may come in the form of machine or sensor logs” (Shine Light on Dark Data – Joe Nicholson via Datameer)
Finally, Lug Bergman of NGDATA wrote this in a sponsored piece in Wired: “It” – dark data – “is different for each organization, but it is essentially data that is not being used to get a 360 degree view of a customer.
Okay, let’s see if we can be a bit more specific about the content of dark data?
Items on the dark data ticket include: Email; Instant messages; documents; Sharepoint content; content of collaboration databases; ZIP files; log files; archived sensor and signal data; archived web content; aged audit trails; operational database backups – full and incremental; roll-back, redo and spooled data files; sunsetted applications (code and documentation); partially developed and then abandoned applications; and, code snippets.
Most importantly, dark data is data that is not actively in use, is underutilised, or is something else. Seriously.
So, the conclusion that some have come to is this: there is a vast collection of data in various formats waiting to be monetised.
Personally, the idea that really grabs my attention is the potential ability to do novel forensic research on email. If only to find out what happened in the past.
For example, maybe it would be fascinating to see how significant challenges were identified, flagged and discussed; how strategic responses to those challenges were formulated, chosen and executed; and, how the outcomes of all of that process were reflected in email communications.
I think that this line of work can be very interesting for some people, and that interesting insights may be uncovered, but I would hate to have to put a tangible value on it, if only to avoid adding to the already galactic magnitudes of nonsense and hype surrounding certain data topics.
There are other more mundane uses of dark data.
Imagine that you are just about to embark on a Data Warehouse project (you really are a late adopter aren’t you), and you want establish a base collection of historical data. Where do you get that historical data from?
Right! Operational databases are not characteristically used to store significant amounts of historical reference data and historical transactions beyond a certain time window; there are performance and other reasons for keeping OLTP systems as lean as possible, so, initial loads of historical data is typically recreated in the Data Warehouse from backups, audit trails or logs.
You don’t need a Chief Data Officer in order to be able to catalogue all your data assets. However, it is still good idea to have a reliable inventory of all your business data, including the euphemistically termed Big Data and dark data.
If you have such an inventory, you will know:
What you have, where it is, where it came from, what it is used in, what qualitative or quantitative value it may have, and how it relates to other data (including metadata) and the business.
What needs to be kept, and for how long, and what can be safely discarded, and when.
The risks associated with the retention or loss of that data.
If you don’t have such a catalogue and have never done a data inventory then a full data inventory and audit seems to be your new best friend.
Simply stated, you may have dark data that has value, or it may be a simple collection of worthless digital nostalgia. But if you don’t know what you have, it may pay to find out what’s there, and if necessary, to let it go.
There is no point in hoarding unneeded and unwanted rubbish data. That is simply not good data management.
Finally a word on all the fuss surrounding dark data.
Failure to monetize when there is value to be obtained from dark data is one thing, claiming that value can be invariably obtained whilst actually not knowing what the data is, or how it could be monetised, is just adding to the mountain of data related ‘nonsense and hype’ doing the rounds these days. Please consider not adding to that mountain.
British Rail, the national UK rail Company, used to be notorious for the number of delays and cancellations to services, and their reasons for failing to meet their obligations became stranger and stranger.
In winter, it would snow and there would be problems. And people would ask ‘how come you couldn’t deal with the snow this year, we’ve had snow for centuries?’ And back came the answers ‘Yes, Sir, but this year it was the wrong type of snow’. In autumn (the fall), it was ‘the wrong types of leaves, and ‘the wrong type of rain’, and in Summer, the ‘wrong type of sunshine’ and so on and so forth.
I hope this will not be the excuse from the Big Data and dark data pundits and punters when the much-vaunted and ‘almost’ guaranteed monetisation isn’t frequently realised.
‘Of course Big Data gives you big dollar benefits, it was just littered with the wrong type of data’ or ‘you just weren’t trying hard enough’.
Many thanks for reading.

Hold this thought: To paraphrase the great Bob Hoffman, just when you think that if the Big Data babblers were to generate one more ounce of bull**** the entire f****** solar system would explode, what do they do? Exceed expectations.
I am a mild mannered person, but if there is one thing that irks me, it is when I hear variations on the theme of “Data Warehousing is Big Data”, “Big data is in many ways an evolution of data warehousing” and “with Big Data you no longer need a Data Warehouse”.
Big Data is not Data Warehousing, it is not the evolution of Data Warehousing and it is not a sensible and coherent alternative to Data Warehousing. No matter what certain vendors will put in their marketing brochures or stick up their noses.
In spite of all of the high-visibility screw-ups that have carried the name of Data Warehousing, even when they were not Data Warehouse projects at all, the definition, strategy, benefits and success stories of data warehousing are known, they are in the public domain and they are tangible.
Data Warehousing is a practical, rational and coherent way of providing information needed for strategic and tactical option-formulation and decision-making.
Data Warehousing is a strategy driven, business oriented and technology based business process.
We stock Data Warehouses with data that, in one way or another, comes from internal and optional external sources, and from structured and optional unstructured data. The process of getting data from a data source to the target Data Warehouse, involves extraction, scrubbing, transformation and loading, ETL for short.
Data Warehousing’s defining characteristics are:
Subject Oriented: Operational databases, such as order processing and payroll databases and ERP databases, are organized around business processes or functional areas. These databases grew out of the applications they served. Thus, the data was relative to the order processing application or the payroll application. Data on a particular subject, such as products or employees, was maintained separately (and usually inconsistently) in a number of different databases. In contrast, a data warehouse is organized around subjects. This subject orientation presents the data in a much easier-to-understand format for end users and non-IT business analysts.
Integrated: Integration of data within a warehouse is accomplished by making the data consistent in format, naming and other aspects. Operational databases, for historic reasons, often have major inconsistencies in data representation. For example, a set of operational databases may represent “male” and “female” by using codes such as “m” and “f”, by “1” and “2”, or by “b” and “g”. Often, the inconsistencies are more complex and subtle. In a Data Warehouse, on the other hand, data is always maintained in a consistent fashion.
Time Variant: Data warehouses are time variant in the sense that they maintain both historical and (nearly) current data. Operational databases, in contrast, contain only the most current, up-to-date data values. Furthermore, they generally maintain this information for no more than a year (and often much less). In contrast, data warehouses contain data that is generally loaded from the operational databases daily, weekly, or monthly, which is then typically maintained for a period of 3 to 10 years. This is a major difference between the two types of environments.
Historical information is of high importance to decision makers, who often want to understand trends and relationships between data. For example, the product manager for a Liquefied Natural Gas soda drink may want to see the relationship between coupon promotions and sales. This is information that is almost impossible – and certainly in most cases not cost effective – to determine with an operational database.

Non-Volatile: Non-volatility means that after the data warehouse is loaded there are no changes, inserts, or deletes performed against the informational database. The Data Warehouse is, of course, first loaded with cleaned, integrated and transformed data that originated in the operational databases.
We build Data Warehouses iteratively, a piece or two at a time, and each iteration is primarily a result of business requirements, and not technological considerations.
Each iteration of a Data Warehouse is well bound and understood – small enough to be deliverable in a short iteration, and large enough to be significant.
Conversely, Big Data is characterised as being about:
Massive volumes: so great are they that mainstream relational products and technologies such as Oracle, DB2 and Teradata just can’t hack it, and
High variety: not only structured data, but also the whole range of digital data, and
High velocity: the speed at which data is generated, transmitted and received.
These are known as the three Vs of Big Data, and they are subject to significant and debilitating contradictions, even amongst the gurus of Big Data (as I have commented elsewhere: Contradictions of Big Data).
From time to time, Big Data pundits slam Data Warehousing for not being able to cope with the Big Data type hacking that they are apparently used to carrying out, but this is a mistake of those who fail to recognise a false Data Warehouse when they see one.
So let’s call these false flag Data Warehouse projects something else, such as Data Doghouses.
“Data Doghouse, meet Pig Data.”
Failed or failing Data Doghouses fail for the same reasons that Big Data projects will frequently fail. Both will almost invariably fail to deliver artefacts on time and to expectations; there will be failures to deliver value or even simply to return a break even in costs versus benefits; and of course, there will be failures to deliver any recognisable insight.
Failure happens in Data Doghousing (and quite possibly in Big Data as well) because there is a lack of coherent and cohesive arguments for embarking on such endeavours in the first place; a lack of real business drivers; and, a lack of sense and sensibility.
There is also a willing tendency to ignore the advice of people who warn against joining in the Big Data hubris. Why do some many ignore the ulterior motives of interested parties who are solely engaged in riding on the faddish Big Data bandwagon to maximise the revenue they can milk off punters? Why do we entertain pundits and charlatans who ‘big up’ Big Data whilst simultaneously cultivating an ignorance of data architecture, data management and business realities?
Some people say that the main difference between Big Data and Data Warehousing is that Big Data is technology, and Data Warehousing is architecture.
Now, whilst I totally respect the views of the father of Data Warehousing himself, I also think that he was being far too kind to the Big Data technology camp. However, of course, that is Bill’s choice.
Let me put it this way, if Oracle gave me the code for Oracle 3, I could add 256 bit support, parallel processing and give it an interface makeover, and it would be 1000 times better than any Big Data technology currently in the market (and that version of Oracle is from about 1983).
Therefore, Data Warehousing has no serious competing paragon. Data Warehousing is a real architecture, it has real process methodologies, it is tried and proven, it has success stories that are no secrets, and these stories include details of data, applications and the names of the companies and people involved, and we can point at tangible benefits realised. It’s clear, it’s simple and it’s transparent.
Just like Big Data, right?
Well, no.
See what I mean?
Therefore, the next time someone says to you that Big Data will replace Data Warehousing or that Data Warehousing is Big Data, or any variations on that sort of ‘stupidity’ theme, you can now tell them to take a hike, in the confidence that you are on the side of reason.
Many thanks for reading.
Aligning Big Data: http://www.linkedin.com/pulse/aligning-big-data-martyn-jones
Big Data and the Analytics Data Store: http://www.linkedin.com/pulse/big-data-analytics-store-martyn-jones
A Modern Manager’s Guide to Big Data:http://www.linkedin.com/pulse/managers-guide-big-data-context-martyn-jones
Accomodating Big Data





Aligning Big Data – Chinese version is thanks to Optimus Prime – published on http://www.36dsj.com/archives/23692
36大数据专稿,原文作者:Martyn Jones 本文由1号店-欧显东编译向36大数据投稿,并授权36大数据独家发布。转载必须获得本站及作者的同意,拒绝任何不标明作者及来源的转载!
引言:
为了带来一些类似的简单性,连贯性和完整性的大数据的辩论,我分享一个普遍信息架构和管理的进化模型。
这是对大数据到一个更通用的体系结构框架的调整和布局,架构集成了数据仓库(DW 2.0),商业智能和统计分析。
这个模型目前称为DW 3.0信息提供框架,简称DW 3.0。
回顾
在以前的一篇比较适用的博客名为“Data Made Simple – Even ‘Big Data‘ ”,里面主要有三个粗略类型的数据:企业运营数据;企业过程数据;以及企业信息数据。如下图:

图1-简要数据模型
简而言之数据的类型可以定义在以下几个:
企业运营数据:这是用于应用程序的数据,支持一个企业的日常运营。
企业过程数据:这是从企业系统是运行的测量和管理收集的数据。
企业信息数据:这主要是数据收集的来自内部和外部的数据源,通常最重要来源是企业运营数据。
这三个底层类型数据是DW 3.0基础。
主体
下面的图展示了DW 3.0总体框架::

图2 -DW3.0信息框架
在这个图中有三个主要元素:数据来源,核心数据仓库和核心数据。
数据来源:这个元素涵盖所有当前的来源,可用的数据的品种和数量用来支持“挑战识别”,“选择定义”的过程和决策,包括统计分析方法和场景法
数据仓库:这是一个DW 2.0模型的演化路径。它扩展了数据仓库的范式不仅包括非结构化和复杂的数据,而且执行的信息和结果来源于统计分析之外的核心数据仓库的场景。
核心统计:这个元素涵盖了核心的统计能力,特别是但不限于对于进化的数据量,数据速度,数据质量和数据的多样性。
这模块的重点是核心统计。也将提及到三者的关系和合并的效果。
核心统计:
下图关注的核心元素模型:

图3 – DW3.0核心统计
上图说明了数据流和信息通过数据采集的过程然后到统计分析和结果的集成。
这个模型还引入了分析数据存储的概念。这可以说是最重要的建筑元素。
数据来源
为了简单起见图中有三个显式指定的数据源(当然依赖的企业数据仓库或数据集市也可以作为一个数据源),但是,我在这篇文章中主要有以下三个数据源:复杂的数据;事件数据;基础数据。
复杂数据:这是结构化或高度复杂的结构化数据文件和其他复杂的数据中包含的文物,如多媒体文件。
事件数据:这是企业过程数据的一个方面,通常在一个细粒度的抽象层次。下面是业务流程日志,互联网web活动日志和其他类似事件数据的来源。这些来源所产生的量往往会高于其他数据源,和那些目前与大数据相关的大量的信息通过追踪即使是最轻微的行为数据覆盖生成一样。例如,有人随意浏览网站。
基础数据:这方面的数据包含可能描述为信号类型数据。通过复杂的事件关联和组件分析产生的连续高速流或者高度动荡的的数据。
革命从这里开始
在这里我将稍微突出建筑元素背后的一些指导原则。
没有业务就没有理由这样做:这是什么意思呢?这意味着每一个重大行动,甚至是高度投机活动,必须有一个有形的和可信的业务支持。就和“奥马哈圣人”,和“圣诞老人”的区别一样清楚。
架构决策都是基于一个完整的和深刻的理解需要实现什么和所有可用的选择:例如,拒绝使用高性能的数据库管理产品必须是有原因的,即使这原因是成本。不应该基于技术意见,如“我不喜欢供应商”如果对Hadoop有感觉,然后使用它,如果对Exasol或Oracle或Teradata有感觉,然后使用它们。那么你一定是一个技术不可知论者,但不是一个有教条的技术论者。
统计和非传统的数据源是完全集成到数据仓库未来架构前景::建设更多的公司仓库,无论是通过行动或遗漏,将导致更大的效率低下,更大的误解和更大的风险。
架构必须连贯,连贯,可用和成本效益:如果没有,有什么意义,对吧?
没有技术,技艺或方法是短板:我们需要能够低成本纳入任何相关现有的新兴技术。
减少早期性和减少频繁性:大量的数据,特别是在高速运转的是存在问题的。减少它们的存储容量,即使我们不能在理论上减少的速度是绝对必要的。我将详细说明这一点区别。
减少早期性,减少频繁性
这里我扩大早期的主题数据减少过滤和聚合,我们可能会产生越来越多的大量的数据,但这并不意味着我们需要囤积所有它为了得到一些价值。
简单的来说这就是将初始数据进行ETL(提取和转换)尽可能靠近数据生成器。这是数据库适配器的概念,但它可以逆转的。
让我们看一个场景。
一个公司想要实施一些投机性分析每天的每一分钟收集的许多互联网网站活动日志数据成,他们运行大量的日志文件分布式平台减少数据映射。
然后他们可以分析结果数据。
面临的问题,与许多网站被黑客,设计师,而不是工程师、建筑师和数据库专家开发,是乱堆着极大的和笨拙的文物,如大量的日志文件的详细钝角和新鲜感添加数据。
我们需要确保这个挑战可以移除吗?
我们需要重新考虑网络日志,然后我们需要重新设计它。
我们需要能够进行语法分析日志数据,以减少产生的大量数据占用严重设计和详细数据。
我们需要的双重选择,能够不断地将数据发送给一个事件设备,可以用来降低数据量在一个事件会话的基础上。
如果我们必须使用日志文件,用许多小日志文件减少大量的日志文件和更多的日志周期减少几个日志周期。我们还必须最大化并行日志的好处。
所以现在,我们得到了日志数据的使用可以通过日志文件、日志文件由一个事件设备(如工具包的一部分分析数据收集适配器)或发送的设备通过消息传递信号点而来。
一旦数据已经传输(传统文件传输/共享或消息)我们可以进入下一个步骤:ET(A)L -提取、转换、分析和负载。
日志文件,我们通常采用ETL(A)但是当然我们不需要ETL中的E即提取,因为这是直接连接。
再次减少ET(AL)是另一种形式的机制,这就是为什么分析方面包括确保得到的数据通过需要的数据,而没有认可价值的垃圾和噪音,会尽早并且经常清理。
分析数据存储
分析数据存储(可以是一个分布式数据存储在某个云)支持统计分析的数据需求。这里的数据组织、结构、集成和丰富的持续波动,偶尔需要统计学家和科学家关注数据挖掘。分析数据存储中的数据可以累计或完全刷新。它可以有一个短寿命或有显著高寿命。
分析数据存储的核心是分析数据。不仅可以用于提供数据统计分析过程,但它也可以用来提供长期持久存储分析结果和场景,和未来的一些分析,因此具有“回馈”的能力。
分析数据存储中的数据和信息也可以使用、来源于数据仓库中存储的数据,它也可能受益于拥有自己的专用数据集市专门为这个目的而设计的。
在分析数据存储的统计分析的结果也可能导致反馈用于调优数据,过滤和浓缩的规则,无论是智能数据分析、复杂事件和歧视适配器或ET(AL)工作。
总结
这一定是非常短暂的对于目前的DW 3.0的标签
模型不寻求定义统计或统计分析是如何应用的,已经做了足够多,但如何适应统计在一个扩展的DW 2.0架构,和几乎不需要想出反动和不合身的问题解决方案,可以解决的更好、更有效的方法通过明智、健全的工程原则和适当的明智的应用方法,技术和技巧。
原文:Aligning Big Data


We’ve been told that Big Data is the greatest thing since sliced bread, and that its major characteristics are massive volumes (so great are they that mainstream relational products and technologies such as Oracle, DB2 and Teradata just can’t hack it), high variety (not only structured data, but also the whole range of digital data), and high velocity (the speed at which data is generated and transmitted). Also, from time to time, much to the chagrin of some Big Data disciples, a whole slew of new identifying Vs are produced, touted and then dismissed (check out my LinkedIn Pulse article on Big Data and the Vs).
So, beware. Things in Big Data may not be as they may seem.
It’s not about bigI have been waging an uphill battle against the nonsensical and unsubstantiated idea that more data is better data, but now this view is getting some additional support, and from some surprising corners.
In a recent blog piece on IBM’s Big Data and Analytics Hub (Big data: Think Smarter, not bigger), Bernard Marr wrote that “the truth is, it isn’t how big your data is, it’s what you do with it that matters!”
Elsewhere, SAS echoed similar sentiments on their web site: “The real issue is not that you are acquiring large amounts of data. It’s what you do with the data that counts.”
Can we call that ‘strike one’ for Big Data Vs?
It’s not about varietyIt is claimed that 20% of digital data is structured, it is based on the problematic suggestion that structured data is uniquely relational. It is also claimed that unstructured data includes CSV files and XML data, and this makes up far more than the 20% of the data generated. But this definition is simply wrong.
If anything, CSV data is structured, and XML data is highly structured, and it’s typically regular ASCII data. So it does not add variety, even though it is not structured in the ways that some people might expect, especially if that someone lacks the required knowledge and experience. Simply stated, CSV data is structured, it’s just that it lacks rich metadata, but that doesn’t make it unstructured.
“But”, I hear you say “what about all the non-textual data such as multi-media, and what about the masses of unstructured textual data?”
Take it from me, most businesses will not be basing their business strategies on the analysis of a glut of selfies, home videos of cute kittens, or the complete works of William Shakespeare or Dan Brown. Almost all business analysis will continue to be carried out on structured data obtained primarily from internal operational systems and external structured data providers.
Strike two! Third time lucky?
It’s not even about velocitySo, if we accept that Big Data isn’t really about the data volumes or data variety that leaves us with velocity, right? Well no, because if it isn’t about record breaking VLDBor significant data variety, then for most commercial businesses the management of data velocity becomes either less of an issue or just is no issue. The fact that some software vendors and IT service suppliers set up this ‘straw man’ argument and then knock it down with the ‘amazing powers’ of their products and services, is quite another matter.
Strike three, and counting.
It’s not about the manageability of Big DataWe have been told and time again that the major difference between a data scientist and professional statistician is that the ‘scientists’ know how to cope very well with massive volumes, varieties and velocities of data. Now it turns out that this is also questionable.
According to Bob Violino writing in Information Management (Messy Big Data Overwhelms Data Scientists – 20 February 2015) “Data scientists see messy, disorganized data as a major hurdle preventing them from doing what they find most interesting in their jobs”. So, when it comes to data quality and structure the ‘scientists’ don’t really have an advantage over professional statisticians.
Last year Thomas C. Redman writing in the Harvard Business Review (Data’s Credibility Problem) noted that when Big Data is unreliable “managers quickly lose faith” and “and fall back on their intuition to make decisions, steer their companies, and implement strategy” and when this happens there is a propensity to reject potentially “important, counterintuitive implications that emerge from big data analyses.”
Strike four?
The new analytics aren’t newData science and Big Data analytics are the new kids on the block, aren’t they?
Well, here are some real life scenarios.
A major banking equipment supplier: A lot of banking equipment is hybrid analogic-digital, a simple example of this would be a photo copier or a physical document processing device. One major supplier decided to incorporate the capture of sensor data produced by their devices to predict failure and problems. Predictive preventive maintenance rules are created and corroborated using the data generated by sensors on each customer device, and these rules then get incorporated into the devices logic.
A major IT vendor: What happens when you create an intersection and convergence between technologies, techniques and method from areas of mainstream IT, data architecture and management, statistics (quantitative and qualitative analytics) and data visualisation, artificial intelligence/machine learning and knowledge management? This is precisely what one of the main European IT vendors did, and the idea proved to be quite attractive to customers, prospects and investors.
A major integrated circuit supplier: The testing of ICs at the ‘fabs’ (manufacturing plants) generates serious amount of data. This data is used to detect errors in the IC manufacturing process, it is captured and analysed in as near real-time as possible, which is necessary due to the costly nature of over-running the production of faulty ICs. To get around this problem the company uses a combination of fast data capture, transformation and loading of data into a data analytics area to ensure early and precise problem detection.
All Big Data Analytics success stories?
The first happened in 1989, the second in 1993 and the third in 2001. Yes, Big Data and Big Data analytics are sort of newish.
Strike five.

What is science?
According to Vasant Dhar of the Stern School of Business (Data Science and Prediction), Jeff Leek (The key word in “Data Science” is not Data, it is Science), and repeated on Wikipedia, “In general terms, data science is the extraction of knowledge from data”. Well, excuse me if I beg to differ. I have seen data scientists at work, and the word science doesn’t actually jump out and grab you. It’s difficult to make the connection, just as it is to accurately connect some popular science magazines with fundamental scientific research.
If a professional and qualified statistician wants to label themselves a data scientist then I have no issue with that, it’s their problem, but I am not willing to lend credibility to the term ‘data scientist’ when it is merely an interesting job title, with at most a tenuous connection to the actual role, and one that is liberally applied, with the almost customary largesse of IT, to creative code hackers and business-averse dabblers in data.
As Hazelcast VP Miko Matsumura suggested in Data Science is Dead “… put “Data Scientist” on your resume. It may get you additional calls from recruiters, and maybe even a spiffy new job, where you’ll be the King or Queen of a rotting whale-carcass of data” and ” Don’t be the data scientist tasked with the crime-scene cleanup of most companies’ “Big Data”—be the developer, programmer, or entrepreneur who can think, code, and create the future.”
Strike six.
And the value is questionableDATA: “Data is a super-class of a modern representation of an arcane symbology.” – Anon
If I had a dollar for every time I heard someone claim that data has intrinsic positive value then I would be as wealthy as Warren Buffet.
If I have said it once, I have said it a hundred time. In order for data to be more than an operational necessity it requires context.
Providing valid data with valid context turns that data into information.
Data can be relevant and data can be irrelevant. That relevance or irrelevance of data may be permanent or temporary, continuous or episodic, qualitative or quantitative.
Some data is meaningless, and there are cases whereby nobody can remember why it was collected or what purpose it serves.
Taking all this into account we can ask the deadly pragmatic question: what value does this data have? Which is sometimes answered with a pertinent ‘no value whatsoever’.
Strike seven.
It is said that Big Data is changing the world, but for all intents and purposes, and shamed by previous Big Data excesses, some people are rapidly changing the definitions and parameters of Big Data, and to position it as being more tangible and down-to-earth, whilst moving it away from its position as an overhyped and dead-ended liability.
Big Data is a dopey term, applied necessarily ambiguously to a surfeit of tenuously connected vagaries, and its time has come and gone. So, let’s drop the Big Data moniker, and embrace the fact that data is data, and long live ‘All Data’, yes, all digital data. Let’s consider all data and for what it’s worth to the business, and not for what some chatterers reckon its value is – having as they do, little or no insight into the businesses to which they refer, or of the data in that these businesses possess.
So, when push comes to shove, is Big Data really about high volumes, high velocity and high variety, or is it in fact about much noise, too much pomposity and abundant similarity leading to unnecessary high anxiety?
Thanks very much for reading.

Big Data is now an inhospitable and unhealthy land inhabited by those who, through accident or design, deceive naïve and sentimental bystanders and those who are willingly mislead.
When all of this Big Data malarkey started it was sort of funny, humorous and occasional witty, especially in the affected, bizarre and the frequently uninhibited ways that freshly-minted self-appointed gurus and experts would “big it up”
Doctor Freud would have had a field day with all of that, being as it was, and for that matter still is, a postmodern mishmash of Riefenstahl, Freddy Mercury and Monty Python on steroids. However, after that extended, operatic and high-camp hiatus it all went downhill.
The Big Data scene is fast becoming an outrageous and brash festival of deception, disinformation and obliviousness. Which is a pity, because it does the industry no good whatsoever.
It is telling that Big Data evangelists, gurus and assorted sycophants cannot even define Big Data adequately, never mind discuss (or for that matter, point at) tangible success stories, without falling into contradictions on all of the key defining characteristics of volume, variety and velocity, and resorting to crude debating devices to avoid or finesse the concerns and the questions.
Almost every morning I check out the industry news, and almost invariably, it comes with new mind-boggling examples of Big Data nonsense.
However, it isn’t always nonsense for nonsense’s sake, there are agendas, there are rational explanations why Big Data has become at the same time, one of the most hyped up fads in the history of IT, and one that its supporters find so difficult to actually explain and justify, in any reasonable sort of way.
Therefore, when it comes to Big Data, beyond the surfeit of platitudes, clichés, bluff and bluster, the only thing in play are the interests of industry, the patrons, the courtesans and their entourage of the innocent and the beguiled.
One of the biggest deceptions in Big Data is in the misleadingly named ‘success stories’. The thing is that most of these success stories that I have ever read have been:
However, it doesn’t stop there.
One of the clearest examples of the questionable nature of Big Data evangelism is when it is used to piggyback Big Data hype on simple, tangible and immediately recognisable artefacts or applications that have little in common with Big Data.
This is an extreme illustration, but it works like this: “iPhones are commercially successful, iPhones are part of Big Data, and therefore Big Data is commercially successful.”
As if the mere conjuring up of association, affinity and proximity will convince people of the great and growing value of Big Data.
What I am also referring to are publicity pieces that may as well have been titled:
Do you recognise similarities?
It’s no big deal, just the use of unreliable, misleading and inappropriate fallacies, dressed up as cute, plausible and accessible collateral. People may think that such things are clever and witty, but they aren’t, it’s just misleading.
Let’s continue with something simple.
Evasion is, in ethics, an act that deceives by stating a true statement that is immaterial or leads to a false deduction. For example, citing events, persons or anecdotes from the history of IT to justify the supposed or imaginary value of Big Data. This is close to the notion of a non sequitur, which of course is an argument, the conclusions from which do not follow from its premise. It falls short of being full-on sophistry, purely because the simplistic, puerile and superficial arguments put forward in favour of Big Data do not match those of the true sophist who seeks to reason with clever but fallacious and deceptive arguments. Too many of the Big Data arguments are fallacious and deceptive, but no one, equipped with a reasonable capacity for critical thinking, should take such ‘arguments’ as valid.
Hold this thought: Big Data hype is a viper’s nest of logical fallacies, white lies and disinformation.
Just when I think things could not get any weirder, they do, and Big Data ceiling of hyperbole rises even higher, up to the rarer atmosphere of extreme tendentiousness.
There is a growing mass of Big Data hoop-la, hyperbole and flim flam that exceeds all previously bounds of overstatement, solecism and confabulation. This is where the real volumes, varieties and velocities are in Big Data; in hokie.
We live, as Oscar Wilde said in his day, in and age of surfaces. Yes, superficiality, puerility and short-termism are the competing orders of the day. However, I am still amazed – and maybe wrongly so – by what ostensibly professional, experienced and knowledgeable people are willing, able and prepared to accept, especially when it comes to Big Data flim flam sauce.
Here are some examples of the nonsense about Big Data that is taken as gospel by ‘adults’:
Data Warehousing is part of Big Data: No comment.
Big Data will replace Enterprise Data Warehousing: People can’t even explain the features and benefits of Big Data. I try it make it as easy as possible, ‘if you can’t say it, point to it’. But, seriously, people can’t even relate tangible and credible Big Data success stories, never mind show how it will replace Enterprise Data Warehousing, whether that’s the Inmon or Kimball flavour, take your pick.
Everyone and every organisation can benefit from Big Data: If people can’t explain this, and they don’t in terms of tangible benefits, then the claim should remain questionable.
Data Scientists will replace Statisticians: Why is that so? It is claimed that Data Scientists are uniquely equipped to handle massive volumes, varieties and velocities of data – well, as it turns out, this isn’t certain either.
Big Data is in its infancy: I think we may be confusing infancy with lack of real traction, and of time and place utility.
You cannot be serious: Just what are people talking about here? I have read vague, naïve and ill-informed pieces about data management, data architecture, data warehousing, reporting, business intelligence and a plethora of etcetera that have been passed off as observations and commentary on Big Data. So, what makes people recycle hackneyed, misleading and badly conceptualised ‘content’?
In the commentary on one of Bernard Marr’s pieces on LinkedIn (a professional networking site) I observed that no one can adequately explain what Big Data is without falling into contradictions and fancies, and no one seems to be capable or willing to provide tangible success stories.
Bernard responded to this comment by pointing out “the reason for that is that Big Data means different things to different people.”
Fair enough. It’s an explanation.
That said, I have always had more than a tenuous dislike of postmodern thinking, in fact most things ‘postmodern’. Call me old fashioned, jaded or cynical, but to me, the idea that everything can mean anything is an aberration that I prefer to leave to others.
I am at a loss to explain why so many reasonable people are willing to embrace the hype surrounding Big Data and Big Data Analytics, including the attendant surfeit of nonsense, incongruences and contradictions, and from my perspective, it defies reason and good sense.
Therefore, I will just end again with a fabulous quote from Ben Goldacre:
“You cannot reason people out of a position that they did not reason themselves into”.
Many thanks for reading.

Please note: This is an edited version of a previous piece with a similar name, but focusing solely on the three main Vs of Big Data.
What we’ve been toldWe’ve been told that business Big Data is the greatest thing since sliced bread, and that its major characteristics are:
Which is a simple and straightforward means of classification. Big Data is about massive volumes, high variety and high velocity. Right?
It’s not about bigI have never bought into the idea that more data is necessarily better data, or that it provides better focus or leads to increased insight, in fact I have been quite vocal with my contrarian opinion, but now this view is getting some additional support, and from some surprising corners.
In a recent blog piece on IBM’s Big Data and Analytics Hub (Big data: Think Smarter, not bigger), Bernard Marr wrote that “the truth is, it isn’t how big your data is, it’s what you do with it that matters!”
Over at Fierce Big Data it was Pam Baker who stated that “the term big data is unfortunate because it’s really not about the size of the data”. (Big data is not about petabytes, but complex computing).
Elsewhere, SAS echoed similar sentiments on their web site: “The real issue is not that you are acquiring large amounts of data. It’s what you do with the data that counts.”
Well, apparently Big Data isn’t about “massive volumes” of data.
Strike 1!

It is claimed that 20% of digital data is structured, it is based on the problematic suggestion that structured data is uniquely relational.
It is also said that unstructured data includes CSV files and XML data, and this makes up far more than the 20% of the data generated. But this definition is wrong.
If anything, CSV data is structured, and XML data is highly structured, and it’s typically regular ASCII data. So there it does not add variety, even though it is not structured in the ways that some someone might expect, especially if that someone lacks the required knowledge and experience. Simply stated, CSV data is structured, it’s just that it lacks rich metadata, but that doesn’t make it unstructured.
“But”, I hear you say “what about all the non-textual data such as multi-media, and what about the masses of unstructured textual data?”
Take it from me, most businesses will not be basing their business strategies on the analysis of a glut of selfies, juvenile twittering, home videos of cute kittens, or the complete works of William Shakespeare. Almost all business analysis (whether done by a professional statistician or a data scientist) will continue to be carried out using structured data obtained primarily from internal operational systems and external structured data providers.
Variety, Sir? No problem.
Strike two!

So, if we accept that Big Data isn’t really about the massive data volumes or high data variety then that leaves us with velocity. Because if it isn’t about record breaking VLDB or significant data variety, then for most commercial businesses the management of data velocity becomes either less of an issue or just is no issue.
Even in some extreme circumstances, one can explore the suggestion that data sampling can remove issues with data volume as well as velocity.
However, the fact that some software vendors and IT service suppliers set up this‘straw man’ velocity argument and then knock it down with the ‘amazing powers’ of their products and services, is quite another matter.
So, is it really about velocity?
Strike three!

Big Data is a dopey term, applied necessarily ambiguously to a surfeit of tenuously connected vagaries, and its time has come and gone. Let’s dump the Big Data moniker, and the 3 Vs along with it, and embrace the fact that data is data, there will always be more of it.
So, let’s consider ‘all data’ and principally for its time and place utility.
If there is something that you are not sure about or have questions with then please leave a comment below or email me.
Thanks very much for reading.

Hold this thought: Big Data is King.
Is there just nothing that Big Data isn’t capable of fixing? From terrorism, world hunger, Ebola, HIV, fraud, money laundering and hiring the ‘right’ people through to winning the lottery, curing hangovers, arranging entrapment and finding the love of your life. Big Data is King. (more…)

Hold this thought: There are real golden nuggets of data that many organisations are oblivious to. But first let’s look at business process management. (more…)

Normally I would send such people to see a specialist – no, not a guru, but a sort of health specialist, but because this has happened to me so many times now, I eventually decided to put pen to paper, push the envelope, open up the kimono, and to record my advice for posterity and the great grandchildren.
So, here are my top seven tips for cashing in quick on the new big thing on the block.

1 – A business opportunity for faith
Like every new religion, trend or fad, Big Data has its own founding myths, theology and liturgy, and there is money to be made in it; loadsa lovely jubbly money. By predicating and evangelising Big Data you will be welcomed with open arms into the Big Data faith, and will receive all the attendant benefits that will miraculously and mysteriously fall upon you and your devout friends. Go on, I dare you. Be a Big Data guru, a shepherd to a flock of sheep, and enjoy the wealth, health and happiness that most surely will come your way. You too can look cool in red Prada slippers, a flattering and flowing gown and matching accessories.

2 – Acquire it, multiply it, weigh it, mark it up and sell it on
Simply stated, this is about acquiring other people’s data, by sacred means or profane, marking it up and then selling it on. The value you add is that you act as a trusted conduit, a conduit for good. You may care to enrich the data, swop the order of data, replicate and embellish data, make stuff up, etc. which all serves to ‘add value’ to the data. You may even consider adding nuggets of value to the data, just for kicks and giggles. My best friend’s favourite is injecting the good old ‘diaper and beer’ and ‘friends and family’ clichés into every Big Data collection, as it never fails to thrill, please and delight.

3 – Anything can be anything
The good thing about making money from Big Data is that it doesn’t need to be anything to do with Big Data. Make a 20GB Enterprise Data Warehouse? Call it a Big Data success. Sell 20 boxes of dodgy doughnuts down the alternative market? Proclaim a Big Data triumph. Sell your digital porn stash to your best mate? Point to the incredible invisible hand of the Big Data market at work. See what I’m doing there. Anything can be anything, and you too can cash in on that opportunity, big time.

4 – Big Data Patronage
Tense, nervous headaches? Do you like making up stories about Big Data, or for that matter anything else? Are you a natural born fibber but are strapped for cash? Then worry no longer. If you get a Big Data patron you will be sorted for ‘life’; get two and you’ll be sorted for the afterlife as well. With a Big Data patron you can get the most tenuous, crappiest and superficial of pieces published, promoted and vaunted – globally. Can’t make it up yourself, then outsource and offshore it, after all, just get the keywords right for SEO ranking and the gullible will flock to you in droves. The down side of this profession is that you will be targeted for writing half-truths, quarter-truths and downright lies, and you will be pilloried as a purveyor of rank hyperbole. But don’t worry, take heart and never lose the faith, you will be in good company. As one Big Data guru was want to say ” If you repeat a lie often enough, people will believe it, and you will even come to believe it yourself.” Amen! brother.

5 – Big Data Certification
By 2016 there will be global demand for 30 billion Big Data professionals. Are you prepared to cash in on that inevitability? No? Then consider this.
One of my best friends makes his living as a completely phony Big Data Scientist. For two hundred bucks he can make you a Data Scientist or a Big Data guru. Some guys give you an education but this guy gives you immediate access to high paying jobs, sex and a life in the city. Moreover, for an extra 250 bucks you can also become a certified Big Data Trainer, which will allow you to do unto others what has been done unto you.

6 – Creative Technology Reuse
Big Data has heralded in the biggest innovations known in the history of computing, and arguably in the entire history of humankind. One of those new inventions has been the now widely acclaimed and revolutionary ‘flat file data base’ (FFDB), and this has been accompanied with developments in low level operating system primitives that allow for the processing of these collections and hierarchies of FFDBs. So, if one has a mind to do so, one can get some real business leverage off of these new tendencies by borrowing 21st century technology found in old operating system hacks from the sixties and seventies and eighties and nineties and… Well, the point is that in order to get serious funding it is no longer good enough to have a half page business plan, it is also necessary to eke out ‘stuff’ that works within the new paradigms of Big Data and Big Data Analytics. For my next venture I will be looking for serious funding for my ‘Arbitrary Dawdle Down Data Street’ (AD3S) Big Data Analytics platform, a platform designed to support virtual 1k bit processing and the massively parallel provision of global regular expression search and match (S&M), concatenation and listing, and cooperative data-driven and streamed data extraction and reporting. I’m hoping to attract the attention of governments, the EU, the Manic Street Preachers, the UN, China, Vladimir Putin, the DOD, HP, Oracle, Gartner, Lana Del Rey, Deloitte and IBM. So, this is going to be absolutely massive. Word!

7 – Big Data Brokerage
According to leading management consultants and industry watchers Gartner, McKinsey and Deloitte, data needs to be managed and accounted like any other asset, such as money. To get into a similar view-point requires a massive leap of faith, but it is a conversion that might drive dividends. One avenue to be explored in eking out value from the apparently massively valuable Big Data lakes, silos and pools is through the operation of a Big Data Brokerage. A Big Data Brokerage is a business whose main responsibility is to be an intermediary that puts Big Data buyers and Big Data sellers together in order to facilitate a transaction. Big Data Brokerage companies are compensated via commission after the Big Data transaction has been successfully completed. They may also charge introductory fees. Just imagine the wealth of business opportunities in that. You could become the Goldman Sachs of data.
I hope you enjoyed this piece and would be pleased to hear your views on this and other subjects.
Whilst I understand the attraction and even the need of creating a new and significant growth industry, I would also advise a degree of restraint, and whilst I see that “Big Data” (the consideration of the potential value of All Data) has its allure, I also think that some good sense and informed caution should also prevail.
Thank you so much for reading.
Martyn Richard Jones

Fueled by the new fashions on the block, principally Big Data, the Internet of Things, and to a lesser extent Cloud computing, there’s a debate quietly taking please over what statistics is and is not, and where it fits in the whole new brave world of data architecture and management. For this piece I would like to put aspects of this discussion into context, by asking what ‘Core Statistics’ means in the context of the DW 3.0 Information Supply Framework.
The following diagram illustrates the overall DW 3.0 framework:
There are three main concepts in this diagram: Data Sources; Core Data Warehousing; and, Core Statistics.
Data Sources: All current sources, varieties, velocities and volumes of data available.
Core Data Warehousing: All required content, including data, information and outcomes derived from statistical analysis.
Core Statistics: This is the body of statistical competence, and the data used by that competence. A key data component of Core Statistics is the Analytics Data Store, which is designed to support the requirements of statisticians.
The focus of this piece is on Core Statistics. It briefly looks at the aspect of demand driven data provisioning for statistical analysis and what ‘statistics’ means in the context of the DW 3.0 framework.
The DW 3.0 Information Supply Framework isn’t primarily about statistics it’s about data supply. However, the provision of adequate, appropriate and timely demand-driven data to statisticians for statistical analysis is very much an integral part of the DW 3.0 philosophy, framework and architecture.
Within DW 3.0 there are a number of key activities and artifacts that support the effective functioning of all associated processes. Here are some examples:
All Data Investigation: An activity centre that carries out research into potential new sources of data and analyses the effectiveness of existing sources of data and its usage. It is also responsible for identifying markets for data owned by the organization.
All Data Brokerage: An activity that focuses on all aspects of matching data demand to data supply, including negotiating supply, service levels and quality agreements with data suppliers and data users. It also deals with contractual and technical arrangements to supply data to corporate subsidiaries and external data customers.
All Data Quality: Much of the requirements for clean and useable data, regardless of data volumes, variety and velocity, have been addressed by methods, tools and techniques developed over the last four decades. Data migration, data conversion, data integration, and data warehousing have all brought about advances in the field of data quality. The All Data Quality function focuses on providing quality in all aspects of information supply, including data quality, data suitability, quality and appropriateness of data structures, and data use.
All Data Catalogue: The creation and maintenance of a catalogue of internal and external sources of data, its provenance, quality, format, etc. It is compiled based on explicit demand and implicit anticipation of demand, and is the result of an active scanning of the ‘data markets’, ‘potential new sources’ of data and existing and emerging data suppliers.
All Data Inventory: This is a subset of the All Data Catalogue. It identifies, describes and quantifies the data in terms of a full range of metadata elements, including provenance, quality, and transformation rules. It encompasses business, management and technical metadata; usage data; and, qualitative and quantitative contribution data.
Of course there are many more activities and artifacts involved in the overall DW 3.0 framework.
Statistics, it is said, is the study of the collection, organization, analysis, interpretation and presentation of data. It deals with all aspects of data, including the planning of data collection in terms of the design of surveys and experiments; learning from data, and of measuring, controlling, and communicating uncertainty; and it provides the navigation essential for controlling the course of scientific and societal advances[i]. It is also about applying statistical thinking and methods to a wide variety of scientific, social, and business endeavors in such areas as astronomy, biology, education, economics, engineering, genetics, marketing, medicine, psychology, public health, sports, among many.
Core Statistics supports micro and macro oriented statistical data, and metadata for syntactical projection (representation-orientation); semantic projection (content-orientation); and, pragmatic projection (purpose-orientation).
The Core Statistics approach provides a full range of data artifacts, logistics and controls to meet an ever growing and varied demand for data to support the statistician, including the areas of data mining and predictive analytics. Moreover, and this is going to be tough for some people to accept, the focus of Core Statistics is on professional statistical analysis of all relevant data of all varieties, volumes and velocities, and not, for example, on the fanciful and unsubstantiated data requirements of amateur ‘analysts’ and ‘scientists’ dedicated to finding causation free correlations and interesting shapes in clouds.
This has been a brief look at the role of DW 3.0 in supplying data to statisticians.
One key aspect of the Core Statistics element of the DW 3.0 framework is that it renders irrelevant the hyperbolic claims that statisticians are not equipped to deal with data variety, volumes and velocity.
Even with the advent of Big Data alchemy is still alchemy, and data analysis is still about statistics.
If you have any questions about this aspect of the framework then please feel free to contact me, or to leave a comment below.
Many thanks for reading.
Catalogue under: #bigdata #technology big data, predictive analytics
[i] Davidian, M. and Louis, T. A., 10.1126/science.1218685

Good morning fellow consumers; here’s a pop quiz question: What does Big Data have in common with Robitussin? Think about, take your time.
Okay, times up!
Robitussin is a legal pharmaceutical product commonly associated with coughs, colds and flu combinations. (more…)

It’s no wonder that truth is stranger than fiction. Fiction has to make sense.
Mark Twain
It’s Friday morning in London’s trendy Canary Wharf, and I have been asked to facilitate a local meeting of the Digital Violence and Dogma Victims Group, the self-help recovery chain for those who have fallen foul of the pernicious and debilitating effects of IT dogma, organisational autism and insider thuggery and blackmail.
There are twelve of us in the old church hall. We sit in a circle, to facilitate communication. After a more formal welcome and brief introduction the floor is opened up for people to talk about whatever they want to talk about. There is silence. This is normal. There are a few new faces.
“Pantxo!” I look across at Pantxo; he is staring out the window at the falling rain. He can usually talk the legs off a giraffe, but today he is having none of it. Sensing that things are not going too well, I enter into my routine of floridly and inanely relating well-worn anecdotes from the distant annals of IT history.
As I am entering my tenth lap of the track of tedium, one of the new members picks up enough courage to chime in, first nervously and then with the increasing confidence of someone who knows exactly what they are talking about and precisely what they are going to say.
“Hello. My name is Crème”. A woman in a blue adidas tracksuit looks around the room.
“Yes. My name is Crème Brûlée; you may well have heard of it from twitter, the tabloids and the TV… oh, and the novel Absolute Beginners… I used to be the CIO of a major household name.”
She pauses and looks into the middle distance, searching for the truth, tip toeing around the pain.
“This is a bit embarrassing – awkward maybe would be a better word – but what I want to unburden upon you all today is the story of how I outsourced my Data Warehouse, my Business Intelligence, my Big Data, my MDM, my CRM, my family and my life”.
She takes a deep breath and continues; making a point of looking at each of her fellow members in turn as she does so.
“About five summers ago, I feeling a bit lost, which was unusual for me, a strange and novel experience, so I decided that I really needed to do something to turbo-charge my career prospects and to get things moving faster in my part of the organisation. I wanted to excel, and I wanted to be seen doing so, by the right people, and recognised as such.”
“In the spring of that year I had been to a management conference with some of our senior IT management team, some of whom are also here today. Okay, I won’t single out any one of you, because you know who you are.”
“As part of the week-long conference we were wined and dined, stroked and cajoled, flattered and sweet talked by a whole entourage of sales execs from the technology and service providers. They were telling us that the future was in outsourcing and offshoring as much as we could, yes even Big Data and Data Warehousing and Analytics, and they were bewitching us with stories of future successes, of IT paradise and professional nirvana. We in turn wanted to believe, needed to believe, desired to believe. All of this was reinforced by the so called independent industry analysts who insisted, in their agnostic way, that we should seize the moment, with courage, determination and illusion.”
“When I got back to the ranch my mind became occupied with other things, but I didn’t entirely forget the compelling messages that I had brought away with me from the conference.”
“Nothing happened for a couple of months until, one day and out of the blue, things came to a head.”
“We had recently acquired a media news and entertainments business – Media Macaroni International, and we were planning on integrating their general ledger into the corporate IT landscape. One morning I received a call from the CIO of the newly acquired company, inviting me to their site for a meet and greet event.”
“So I moseyed on down to Tinsel Town and got a briefing from not only the CIO, but the full board of directors of Media Macaroni, the ‘hasta la pasta’ of Big Data Analytics ad-hoc performance alignment.”
“To cut a long story short, they basically put me on the spot. Either I integrate the entire Data Warehousing, MIS, Big Data, Analytics and MDM across the expanded corporate body in 9 to 15 months, or we would have serious problems of convergence and market credibility. The message couldn’t have been clearer. Either I got our act together and made this acquisition work, or what looked like a humongous hot potato could land in my lap anytime soon. It was a career changing risk that I needed to address.”
“I told the directors there and then that the mission was going to be incredibly difficult to fulfil. However, the mood quickly changed.”
“Their CIO looks across the table and tells me that he can help me out of my hole. My hole? What the freak! You see, we have employed a service company that does most of our IT work for us, and according to us at Media Macaroni they are simply the bee’s knees, the best thing since chopped liver on rye, the biz.”
“So ‘who are these guys’, I ask. And after a brief hiatus that seemed to last forever, back came the ominous response: The Taffia Connection.”
To be continued…
Many thanks for reading.
Channel: #IT #BigData
As always, please share your questions, views and criticisms on this piece using the comment box below. I frequently write about strategy, organisational leadership and information technology topics, trends and tendencies. You are more than welcome to keep up with my posts by clicking the ‘Follow’ link and perhaps even send me aLinkedIn invite. Also feel free to connect via Twitter, Facebook and the Cambriano Energy website.
File under: Good Strat, Good Strategy, Martyn Richard Jones, Martyn Jones, Cambriano Energy, Iniciativa Consulting, Iniciativa para Data Warehouse, Tiki Taka Pro